devtools.mcpEval
devtools: devtools.mcpEval ()
devtools: devtools.mcpEval ("Sandbox")
Serve the Model Context Protocol on standard input and output, with evaluation.
devtools.mcpEval () is devtools.mcp plus the two tools that run
code, octave_eval and octave_test.
It reads newline-delimited JSON-RPC messages from standard input, answers
each one, writes the answer to standard output, and returns only when
standard input reaches end of file.
This is a separate entry point on purpose. It is a separate
function, a separate launch command and a separate entry in a host’s
configuration, so that a host configured for devtools.mcp cannot reach
these tools however a model asks. The read-only server evaluates no code,
runs no user function and writes nothing, which is what lets a user grant it
blanket permission; this one does all three, and is meant to be configured
under its own name, conventionally "octave-eval", so that the
permission rules for the two can differ.
Configuration
Configure a host to launch it under its own name with:
octave-cli -q --no-init-file --eval "pkg load devtools; devtools.mcpEval ()" |
The rules are those of devtools.mcp and they matter as much here: do
not shorten the command, name octave-cli.exe in full on Windows, and
load whatever packages this server is meant to see, since it sees the ones
its own launch command loads and no others.
Tools
The five read-only tools of devtools.mcp are served here as well,
plus two that run code:
octave_evaloctave_testTesting
octave_test runs the built-in tests of one function or file and
reports how many passed, with the assertion behind each failure. It
resolves a name through which and then runs the file, which is
what lets it test a namespaced function or a class method: core’s
test cannot resolve devtools.jsonrpcError and answers
"does not exist in path" with a count of zero, and a count of zero
reads as "no tests" rather than as a name it could not resolve.
It takes no workspace. Tests run in a context of their own every time, so that what passed cannot depend on what was evaluated before.
Workspaces
Evaluation state is held in a workspace, named by an opaque handle.
Every call names one: "new" opens a workspace and the reply gives
its handle, and that handle passed back continues it. This is what the
protocol requires, state that spans requests being referenced by an explicit
identifier rather than by the connection it arrived on. A handle lives
until the process ends or until it is the oldest of more than eight, and a
call naming one that is gone is a tool error that says so rather than a
fresh workspace that says nothing.
The handle is required rather than optional, and that is a deliberate
departure from the specification’s suggested shape of a separate creation
tool: a model that omits an argument is the measured case, and an omitted
handle read as "start clean" would lose a workspace in silence.
What contains it, and what does not
Output is captured twice over, and the two halves catch different things.
evalc takes every route to standard output that stays inside the
interpreter, and __devtools_capture__ holds descriptor 1 over a file
for the length of the call, which is the only thing that catches a
subprocess: a child inherits the descriptor and writes past
evalc entirely, into the stream that carries the protocol. What a
child printed comes back labelled in the reply.
Because that containment is at the descriptor and not at a name,
builtin ("system", …) does not get around it. Where the
package was installed without a compiler and __devtools_capture__
could not be built, the containment moves to the two forms that let a child
inherit descriptor 1: system is shadowed by one that asks for the
output back and prints it through the interpreter, where evalc takes
it, and popen by one that refuses its write mode, whose output
nothing there could read. That is weaker in one way, since a shadow at a
name is defeated by builtin, and it costs an asynchronous
system, whose output core will not return at all.
input and keyboard are shadowed in either case, since there is
no terminal for them to read from.
The deadline
An evaluation that is still running after twenty seconds is stopped, and the reply says so. What the code assigned before it was stopped is still in the workspace and what it printed is still returned, which is the difference between this and letting the host restart the process.
The stopping is the interpreter’s own interrupt, the mechanism Ctrl-C uses,
raised from a thread and caught in __devtools_guard__. It cannot be
done in Octave: measured on 11.2.0, an interrupt raised this way unwinds
straight through try and takes the process with it, honouring
unwind_protect on the way but never being caught.
Set DEVTOOLS_EVAL_SECONDS in the launch command for an installation
whose work honestly takes longer, up to six hundred. It is not a tool
argument, since that would cost tokens in every request and is a decision for
whoever configures the server rather than for the model.
The deadline does not recover everything. A call wedged inside one long native call, or blocked on a read, reaches no checkpoint at which the interrupt can be noticed, and the host restarting the process is what remains. Where the oct-files could not be built there is no deadline at all.
Sandbox
devtools.mcpEval ("Sandbox") serves a program rather than a model,
inside a sandbox where this machine can build one: on GNU/Linux
with bwrap from the bubblewrap package and
prlimit installed, and on macOS with the system’s own
sandbox-exec. Where it cannot, it serves without one and says
so. A host that does not read the report below should not be given this
option on a machine without a sandbox.
The sandbox is settled when the server starts and does not change while it
runs. It is stated in _meta["io.github.pr0m1th3as.devtools/sandbox"]
in the reply to initialize, and, under the stateless protocol
revision, which has no initialize, in every reply; the
instructions say the same to a model. It is one of three states:
"active""failed""unavailable"bwrap,
prlimit or sandbox-exec missing. The server serves
unconfined. For the last two, _meta["io.github.pr0m1th3as.devtools/sandboxReason"]
says why, and standard error logs it as the server starts.
Either way it offers octave_call and octave_test beside the
read-only tools, and not octave_eval: every call starts from the same
state in a process of its own, so a workspace would carry nothing.
octave_call is for programs, which read its structured result, its
text being a summary without the values. It runs no code text: it calls
one function by name on typed arguments, a matrix as a list of rows, a flat
list being a column, and a range carrying each cell’s kind and value, and
returns each output as typed cells row by row, dates as serial numbers from
the document’s null date, with anything the function printed beside them.
A call that crashes the interpreter comes back as an error, and the server
keeps serving.
The folders it may read and the packages it loads are set in the launch
environment, never in the command. DEVTOOLS_SANDBOX_FOLDERS holds
absolute folder paths separated by pathsep, and
DEVTOOLS_SANDBOX_PACKAGES holds package names separated by commas,
loaded in that order. Nothing checks whether two of them conflict. The
folders are on the load path, ahead of the packages, with or without the
sandbox. A host launches it like the plain server:
octave-cli -q --no-init-file \
--eval "pkg load devtools; devtools.mcpEval ('Sandbox')"
|
Inside the sandbox
The server first builds the sandbox once, checks it from inside and leaves
it, and serves unconfined, "failed", if that check does not pass.
Only then does it replace its own process with a sandboxed
octave-cli built by devtools.sandboxCommand. The process, its
standard streams and its exit code carry through unchanged. The trial is
there because a replaced process cannot come back: a sandbox found wanting
from inside has no unconfined process left to serve from. Should the check
pass on the trial and fail on the real run, the server stays up, reports
"failed" and offers no tools at all, since serving half-confined is
what the check exists to prevent.
The check is that there is no /usr/bin, no network interface
besides the loopback, an address-space limit in force, and nothing under
/home or the home directory that was not mounted. The limit is this
process’s size plus 2 GB, or plus the number of gigabytes in
DEVTOOLS_SANDBOX_MEMORY, and an allocation beyond it fails with
Octave’s own out-of-memory error. /tmp, whose files are memory too,
holds at most 2 GB, or the number of gigabytes in
DEVTOOLS_SANDBOX_TMP. It lists only the packages that are mounted,
so loading any other says it is not installed. See
devtools.sandboxCommand for what is mounted and what is refused.
Each call runs in a process forked for it, which is killed when it returns
or when the deadline passes, together with every process it started, and
/tmp is emptied before the next call, so that nothing one call does
reaches another.
The two systems confine by different means, so each names the guarantees it
holds rather than claiming the other’s. On macOS a sandboxed server writes
one line to standard error as it starts,
OMP: Warning #179: Function Can't set size of /tmp file failed:, and
then serves normally. OpenMP registers itself by making
/tmp/__KMP_REGISTERED_LIB_<pid>, naming /tmp rather
than reading TMPDIR, and the sandbox does not grant the host’s
/tmp. What that registration detects is a second OpenMP runtime in
the same process, which Octave does not have, so nothing is lost by it. It
is left alone deliberately: granting the write would put a file of the
server’s outside the sandbox at every launch, and suppressing the warning
would hide the next thing to go wrong on the same path.
Without the sandbox
Unconfined, each call still runs in a process of its own and is stopped at
the deadline: forked where the system can fork, and on Windows, which
cannot, an octave-cli started for the call inside a job object by
__devtools_spawn__, which a working compiler builds at installation.
A process a call starts outlives it where it was forked, and the call can
read and write files and use the network as any Octave code can.
What this is not
Unless started with "Sandbox" and reporting it "active",
none of this is a sandbox. Evaluated code can read and write files, use the
network and consume memory exactly as any code in this interpreter can.
Configure this server only where that is acceptable.
See also: devtools.mcp, devtools.selftest, devtools.sandboxCommand
Source Code: devtools.mcpEval