orthonym.jvm_budget#

Note

Internal API. Names and behaviour may change between releases.

Machine-wide JVM concurrency budget, shared by every session and script.

Replaces the blanket folk rule “never run two OPSIN jobs concurrently” with a measured budget and real enforcement. The rule it replaces was never measured; worse, it did not describe the code – eval/harness.py has always run mp.Pool(cpu_count - 4) = 12 concurrent JPype JVMs on this host. Measurements and citations: internal notes.

Why a budget at all, given memory is not the constraint#

Measured on the 16-vCPU / 60 GB host: one OPSIN JVM peaks at ~200-400 MB RSS (the 15.8 GB MaxHeapSize default is reserved address space, not committed), and 32 concurrent JVMs summed to 4.7 GB. Memory would allow 100+.

CPU is the binding constraint, and oversubscription is a pure loss. JVM startup is 0.938 s against a marginal 0.92 ms per name – one launch costs about 1,020 names of real work – so splitting a fixed workload across JVMs re-pays startup per JVM and the startups contend. Throughput falls monotonically: 712 names/s at c=1 down to 166 at c=24. The budget exists to stop unrelated jobs from driving each other down that curve, not to prevent a correctness failure.

Why flock rather than a PID file or pgrep#

The kernel releases an flock when the holding process dies, however it dies – including SIGKILL. So a slot cannot go stale, which is exactly the failure mode a PID file or sentinel has. pgrep additionally cannot express “how many”, and it self-matches (a documented hazard here: an unescaped pattern matches the checking command itself).

This budget is advisory and fail-open by design. It exists to schedule work, never to lose it: if the lock directory is unusable the call proceeds with a warning rather than failing a caller’s run. Set ORTHONYM_JVM_BUDGET=off to disable entirely.

Usage#

from orthonym.jvm_budget import jvm_slots

# a gate: in-process, one JVM with jvm_slots(1, purpose=”v22_gate”):

run_conformance(…)

# an eval harness: one JVM per pool worker with jvm_slots(jobs, purpose=”harness:a dev split”):

pool.map(…)

# see what holds the budget right now (replaces pgrep -f v22_gate) from orthonym.jvm_budget import status for slot in status[“held”]:

print(slot[“purpose”], slot[“pid”])

exception orthonym.jvm_budget.BudgetTimeout#

Bases: RuntimeError

Raised by:func:jvm_slots when on_timeout='raise' and the wait expired.

orthonym.jvm_budget.is_enabled()#
orthonym.jvm_budget.total_slots()#

Total concurrent JVM slots for this machine.

orthonym.jvm_budget.slot_dir()#

Directory holding the lock files. Shared across sessions and worktrees.

Deliberately NOT inside the repo: the budget is a property of the machine, and two worktrees or two sessions of the same checkout must contend over the same slots.

orthonym.jvm_budget.jvm_slots(count=1, *, purpose='unnamed', timeout=None, on_timeout='proceed')#

Reserve count concurrent-JVM slots for the duration of the block.

Parameters:
  • count (int) – JVMs this job will have alive at once. One in-process JPype JVM or one java -jar subprocess is 1; an mp.Pool(n) whose workers each start a JVM is n.

  • purpose (str) – short label recorded in the slot file and shown by status() – e.g. "v22_gate", "harness:a dev split".

  • timeout (float | None) – seconds to wait. None waits indefinitely.

  • on_timeout (str) – "proceed" (default) runs anyway with a warning – correct for a long job a caller must not lose; "raise" raises BudgetTimeout, for callers that would rather refuse.

Never raises on a missing/unwritable lock directory: the budget is advisory and must not be able to fail a naming run.

orthonym.jvm_budget.status()#

Report which slots are currently held, and by what.

The supported replacement for pgrep -f "scripts/v22_gate\.py": it counts, it names the holder, and it cannot self-match the checking command.

orthonym.jvm_budget.main(argv=None)#

python -m orthonym.jvm_budget – show the current budget.