Categories &

Functions List

Function Reference: devtools.mcpEval

devtools: devtools.mcpEval ()

devtools: devtools.mcpEval ("Sandbox")

Serve the Model Context Protocol on standard input and output, with evaluation.

devtools.mcpEval () is devtools.mcp plus the two tools that run code, octave_eval and octave_test. It reads newline-delimited JSON-RPC messages from standard input, answers each one, writes the answer to standard output, and returns only when standard input reaches end of file.

This is a separate entry point on purpose. It is a separate function, a separate launch command and a separate entry in a host’s configuration, so that a host configured for devtools.mcp cannot reach these tools however a model asks. The read-only server evaluates no code, runs no user function and writes nothing, which is what lets a user grant it blanket permission; this one does all three, and is meant to be configured under its own name, conventionally "octave-eval", so that the permission rules for the two can differ.

Configuration

Configure a host to launch it under its own name with:

 
 octave-cli -q --no-init-file --eval "pkg load devtools; devtools.mcpEval ()"

The rules are those of devtools.mcp and they matter as much here: do not shorten the command, name octave-cli.exe in full on Windows, and load whatever packages this server is meant to see, since it sees the ones its own launch command loads and no others.

Tools

The five read-only tools of devtools.mcp are served here as well, plus two that run code:

octave_eval
what running some Octave code produces, in a workspace that persists between calls.
octave_test
how many of a function’s built-in tests pass, and what failed.

Testing

octave_test runs the built-in tests of one function or file and reports how many passed, with the assertion behind each failure. It resolves a name through which and then runs the file, which is what lets it test a namespaced function or a class method: core’s test cannot resolve devtools.jsonrpcError and answers "does not exist in path" with a count of zero, and a count of zero reads as "no tests" rather than as a name it could not resolve.

It takes no workspace. Tests run in a context of their own every time, so that what passed cannot depend on what was evaluated before.

Workspaces

Evaluation state is held in a workspace, named by an opaque handle. Every call names one: "new" opens a workspace and the reply gives its handle, and that handle passed back continues it. This is what the protocol requires, state that spans requests being referenced by an explicit identifier rather than by the connection it arrived on. A handle lives until the process ends or until it is the oldest of more than eight, and a call naming one that is gone is a tool error that says so rather than a fresh workspace that says nothing.

The handle is required rather than optional, and that is a deliberate departure from the specification’s suggested shape of a separate creation tool: a model that omits an argument is the measured case, and an omitted handle read as "start clean" would lose a workspace in silence.

What contains it, and what does not

Output is captured twice over, and the two halves catch different things. evalc takes every route to standard output that stays inside the interpreter, and __devtools_capture__ holds descriptor 1 over a file for the length of the call, which is the only thing that catches a subprocess: a child inherits the descriptor and writes past evalc entirely, into the stream that carries the protocol. What a child printed comes back labelled in the reply.

Because that containment is at the descriptor and not at a name, builtin ("system", …) does not get around it. Where the package was installed without a compiler and __devtools_capture__ could not be built, the containment moves to the two forms that let a child inherit descriptor 1: system is shadowed by one that asks for the output back and prints it through the interpreter, where evalc takes it, and popen by one that refuses its write mode, whose output nothing there could read. That is weaker in one way, since a shadow at a name is defeated by builtin, and it costs an asynchronous system, whose output core will not return at all.

input and keyboard are shadowed in either case, since there is no terminal for them to read from.

The deadline

An evaluation that is still running after twenty seconds is stopped, and the reply says so. What the code assigned before it was stopped is still in the workspace and what it printed is still returned, which is the difference between this and letting the host restart the process.

The stopping is the interpreter’s own interrupt, the mechanism Ctrl-C uses, raised from a thread and caught in __devtools_guard__. It cannot be done in Octave: measured on 11.2.0, an interrupt raised this way unwinds straight through try and takes the process with it, honouring unwind_protect on the way but never being caught.

Set DEVTOOLS_EVAL_SECONDS in the launch command for an installation whose work honestly takes longer, up to six hundred. It is not a tool argument, since that would cost tokens in every request and is a decision for whoever configures the server rather than for the model.

The deadline does not recover everything. A call wedged inside one long native call, or blocked on a read, reaches no checkpoint at which the interrupt can be noticed, and the host restarting the process is what remains. Where the oct-files could not be built there is no deadline at all.

Sandbox

devtools.mcpEval ("Sandbox") serves a program rather than a model, inside a sandbox where this machine can build one: on GNU/Linux with bwrap from the bubblewrap package and prlimit installed, and on macOS with the system’s own sandbox-exec. Where it cannot, it serves without one and says so. A host that does not read the report below should not be given this option on a machine without a sandbox.

The sandbox is settled when the server starts and does not change while it runs. It is stated in _meta["io.github.pr0m1th3as.devtools/sandbox"] in the reply to initialize, and, under the stateless protocol revision, which has no initialize, in every reply; the instructions say the same to a model. It is one of three states:

"active"
The sandbox was built and verified from inside, and every call runs in it.
"failed"
The machine has the mechanism, but the sandbox did not start or did not pass its check. The server serves unconfined.
"unavailable"
The machine has no mechanism for one: another system, or bwrap, prlimit or sandbox-exec missing. The server serves unconfined.

For the last two, _meta["io.github.pr0m1th3as.devtools/sandboxReason"] says why, and standard error logs it as the server starts.

Either way it offers octave_call and octave_test beside the read-only tools, and not octave_eval: every call starts from the same state in a process of its own, so a workspace would carry nothing. octave_call is for programs, which read its structured result, its text being a summary without the values. It runs no code text: it calls one function by name on typed arguments, a matrix as a list of rows, a flat list being a column, and a range carrying each cell’s kind and value, and returns each output as typed cells row by row, dates as serial numbers from the document’s null date, with anything the function printed beside them. A call that crashes the interpreter comes back as an error, and the server keeps serving.

The folders it may read and the packages it loads are set in the launch environment, never in the command. DEVTOOLS_SANDBOX_FOLDERS holds absolute folder paths separated by pathsep, and DEVTOOLS_SANDBOX_PACKAGES holds package names separated by commas, loaded in that order. Nothing checks whether two of them conflict. The folders are on the load path, ahead of the packages, with or without the sandbox. A host launches it like the plain server:

 
 octave-cli -q --no-init-file \
   --eval "pkg load devtools; devtools.mcpEval ('Sandbox')"

Inside the sandbox

The server first builds the sandbox once, checks it from inside and leaves it, and serves unconfined, "failed", if that check does not pass. Only then does it replace its own process with a sandboxed octave-cli built by devtools.sandboxCommand. The process, its standard streams and its exit code carry through unchanged. The trial is there because a replaced process cannot come back: a sandbox found wanting from inside has no unconfined process left to serve from. Should the check pass on the trial and fail on the real run, the server stays up, reports "failed" and offers no tools at all, since serving half-confined is what the check exists to prevent.

The check is that there is no /usr/bin, no network interface besides the loopback, an address-space limit in force, and nothing under /home or the home directory that was not mounted. The limit is this process’s size plus 2 GB, or plus the number of gigabytes in DEVTOOLS_SANDBOX_MEMORY, and an allocation beyond it fails with Octave’s own out-of-memory error. /tmp, whose files are memory too, holds at most 2 GB, or the number of gigabytes in DEVTOOLS_SANDBOX_TMP. It lists only the packages that are mounted, so loading any other says it is not installed. See devtools.sandboxCommand for what is mounted and what is refused. Each call runs in a process forked for it, which is killed when it returns or when the deadline passes, together with every process it started, and /tmp is emptied before the next call, so that nothing one call does reaches another.

The two systems confine by different means, so each names the guarantees it holds rather than claiming the other’s. On macOS a sandboxed server writes one line to standard error as it starts, OMP: Warning #179: Function Can't set size of /tmp file failed:, and then serves normally. OpenMP registers itself by making /tmp/__KMP_REGISTERED_LIB_<pid>, naming /tmp rather than reading TMPDIR, and the sandbox does not grant the host’s /tmp. What that registration detects is a second OpenMP runtime in the same process, which Octave does not have, so nothing is lost by it. It is left alone deliberately: granting the write would put a file of the server’s outside the sandbox at every launch, and suppressing the warning would hide the next thing to go wrong on the same path.

Without the sandbox

Unconfined, each call still runs in a process of its own and is stopped at the deadline: forked where the system can fork, and on Windows, which cannot, an octave-cli started for the call inside a job object by __devtools_spawn__, which a working compiler builds at installation. A process a call starts outlives it where it was forked, and the call can read and write files and use the network as any Octave code can.

What this is not

Unless started with "Sandbox" and reporting it "active", none of this is a sandbox. Evaluated code can read and write files, use the network and consume memory exactly as any code in this interpreter can. Configure this server only where that is acceptable.

See also: devtools.mcp, devtools.selftest, devtools.sandboxCommand

Source Code: devtools.mcpEval