Skip to main content

The scheduler

classcad-scheduler puts one port in front of a pool of ClassCAD engines on a server. The PDM's workers and editor talk to that port, and the scheduler decides which engine does the work. It starts engines, watches them, restarts the ones that crash and retires the ones nobody needs.

In development

The scheduler is part of the ClassCAD runtime and in development alongside the PDM (branch dev/scheduler). It runs on Windows and Linux, on one server. It isn't in a ClassCAD release yet.

PDM job workers ── HTTP ──────┐                       ┌──▶ w1  127.0.0.1:9094
├──▶ scheduler :9091 ───┼──▶ w2 127.0.0.1:9095
PDM editor ────── WebSocket ──┘ └──▶ h1 heavy work (optional)

Start it​

The scheduler sits next to classcad-cli. Tell it how many engines to start with:

./classcad-scheduler --workers 2
classcad-scheduler <version> listening on 0.0.0.0:9091
clients: http://localhost:9091/api | ws://localhost:9091
status: http://localhost:9091/scheduler/status | /scheduler/workers | /scheduler/sessions

It reads scheduler.ini from the path given with --ini, from where it's started, or from its own folder. The engines get the usual .classcad.ini. Each engine logs to its own file, logs/worker-w1.log and so on. On Windows it's classcad-scheduler.exe. There's also a Docker image that runs the scheduler and its engines on one server; it publishes port 9091 only.

Connect the PDM​

CC_URL=http://127.0.0.1:9091/api      # job workers
EDITOR_CC_URL=wss://pdm.example.com/classcad/ # the editor, through your proxy, when it runs on the server

Nothing else changes. Sessions keep working the way they did against a single engine.

How it decides​

Sessions stay put. A request with a ClassCAD-Session-Id header goes to the engine that holds the session. Generator runs work this way. A WebSocket connection is its own session and is relayed to one engine for its whole life, which is how the editor works. A session is bound when it's created (GET /session) and let go with DELETE /session.

Everything else goes where there's room. A request without a session, like the workers' jobs, goes to a ready engine that's under its memory limit and has fewer than 50 requests in flight. The one with the fewest requests in flight wins, then the one with the fewest sessions, then the one using the least memory. A new session also needs a free slot: 10 sessions per engine by default. If no engine qualifies, the request gets 503 with Retry-After: 2, and the scheduler starts an engine right away.

Heavy work is kept apart. Requests whose ClassCAD calls match the heavy patterns (imports, exports, healing), requests bigger than 10 MB, and requests marked X-ClassCAD-Priority: heavy go to heavy engines, if you run any (heavy_workers). If none is free, they go to the regular engine with the most memory to spare.

The pool grows and shrinks. Every 30 seconds, if no ready engine is idle with free slots, the scheduler starts one, up to max. An engine with no sessions and no work for 10 minutes is retired, down to min. It saves its sessions before it goes.

Crashes. Every engine is asked for its status every 2 seconds. One that misses three answers is taken out and started again, waiting 1 s at first and up to 30 s if it keeps failing. Its sessions are marked lost. A session that was saved to disk (an engine saves sessions that sit idle, and all of them when it shuts down) comes back on another engine with its next request. Any other session gets 410 Gone, and the PDM retries the job. A WebSocket session ends with its engine: the editor has to open the draft again.

Configuration​

scheduler.ini uses the same format as .classcad.ini. Every key is optional.

KeyDefaultWhat it does
[listen] host, port0.0.0.0, 9091Where clients connect
[workers] exe./classcad-cliThe engine
[workers] ini./.classcad.iniThe engines' own configuration
[workers] min, max1, 6Size of the pool
[workers] port_base9094The first engine's port; engines listen on 127.0.0.1 only
[workers] max_sessions10Sessions per engine, 0 for no limit
[workers] session_state_dir./session_stateWhere engines save sessions; shared by all of them
[workers] log_dir./logsOne log per engine
[workers] poll_interval_ms, fail_threshold2000, 3Health checks
[workers] shutdown_grace_ms10000How long an engine gets to save and stop
[workers] heavy_workers, heavy_max_sessions0, 3Engines for heavy work
[workers] static, static_heavyEngines you start yourself, as host:port lists. The scheduler then starts, scales and restarts nothing
[placement] memory_threshold_mb3000No new work for an engine above this; heavy_memory_threshold_mb is 8000
[placement] max_inflight_per_worker50Requests in flight per engine
[placement] allow_heavy_fallback1Heavy work may use a regular engine when no heavy one is free
[scaling] enabled, interval_ms, idle_timeout_ms1, 30000, 600000Growing and shrinking the pool
[classification] heavyAPI patterns that count as heavy, like v1/io/import*, v1/io/export*, v1/healing/*, v1/mesh/generate*
[classification] heavy_body_bytes10485760Requests bigger than this count as heavy
[auth] webhook_authA URL that approves each request, see below

Set the heavy patterns yourself: the list is empty unless you do.

Command line: --ini, --workers N, --worker-exe, --worker-ini, --static-workers, --port, --check (prints the configuration it would use, as JSON), --version, --help.

Watching it​

RouteWhat you get
GET /statusThe pool as one: OK, DEGRADED or DOWN, sessions, queue. 503 when no engine is ready
GET /scheduler/statusEngines, how many are ready or down, sessions, requests queued and in flight
GET /scheduler/workersEach engine: state, memory, queue, requests in flight, sessions
GET /scheduler/sessionsEach session and the engine that holds it
GET /scheduler/configThe configuration in use
GET /metricsPrometheus metrics, classcad_scheduler_*

These routes have no access control. Keep port 9091 inside your network and don't forward /scheduler/* or /metrics through your proxy.

Security​

With [auth] webhook_auth set, the scheduler asks that URL before every HTTP request and every new WebSocket: it posts the path and the headers, and a 200 lets the request through. The webhook can name the user in an X-Verified-User response header, which the scheduler passes on to the engine. The webhook doesn't see query strings, and browsers can't add headers to a WebSocket, so until the PDM has its sign-in, an editor connecting through the scheduler belongs on a trusted network.

Good to know​

  • One scheduler per server. Several servers, failover and TLS in the scheduler itself are planned, not built.
  • /status through the scheduler reports the scheduler's version, not an engine's.
  • The scheduler skips ports that are taken, so if you run an engine by hand on 9094, the scheduler's first engine gets the next free port.
  • On stop (Ctrl C, SIGTERM), the scheduler tells every engine to save its sessions and shut down, then waits for them.