The Model Boundary Is an Untrusted Interface

The bot runs a handful of scheduled jobs that use a language model: summarizing a session, reviewing logs, writing reflections. Until recently each one ran through an agentic command-line tool. The job handed over a prompt, the agent went off and read files, ran commands and looked around, and eventually it wrote something back.

That worked. It was also very hard to reason about. Over the last few weeks I moved these jobs onto a plain API planner. Now the job builds a request, sends it through a gateway, and gets text back. In these jobs the model can't browse, run anything, or decide to look at one more file.

I expected a simple swap. What I actually ended up doing was writing down three contracts the old setup had been relying on without ever saying so.

What changed when the model stopped exploring

With the agentic tool, I had never defined the model's inputs. It found its own context, so if a log job needed to see something, it went and got it. I couldn't say what it had looked at, and when it reported "nothing unusual" I couldn't tell whether it had checked everything or only the first thing it opened.

A plain API call has no exploration step. Everything the model sees has to be put into the request by my code. That meant I had to decide what goes in, and once the code was making that decision, it was also responsible for getting it right.

I made the migration reversible. Each job got its own switch between the old and new backend, and the old path stayed in place. I moved jobs one at a time on staging and checked each one's output against what it was supposed to produce. The second-opinion reviewer rejected the first version of the switch outright, so one job went back to the old path until its gaps were closed. The last job to move was the log analyzer, because it depended most on being able to explore.

Contract one: what the model is allowed to see

The log analyzer used to read logs directly. Now the host builds a bounded snapshot: a fixed window of log material, capped in size, put together before the model is called.

A bounded snapshot creates a new problem. If the snapshot was truncated, or a source was missing, or a window came back empty, nothing in the snapshot itself says so. Without the host telling it, the model can't reliably tell whether it's seeing the whole picture, and it may produce a report that sounds more reassuring than the input justifies. "No errors found" from a model that was shown half the logs is worse than no report at all.

So the host now attests to coverage. Along with the snapshot, it records which sources were included, whether anything was cut off, and whether the window was complete. That record comes from the host, not the model. If coverage is incomplete, the report is still delivered, but with a warning attached, so a quiet report never passes for a complete one.

The rule I took from this: when you decide what the model sees, you also have to say how much of the picture that was.

Contract two: what the model is allowed to ask for

The jobs I migrated don't call tools at all. Some of the bot's other jobs still do. In those, the model doesn't run anything itself. It asks for a tool call, and the host carries it out through a client layer. A few of those tools take optional flags, including ones that turn off a fallback path. Fallbacks are a safety feature. They exist so that a failed primary path doesn't quietly turn into "no data."

I found that the model could set those flags itself. Nothing malicious was going on. The flags were in the tool schema, and the model filled them in because it could. But that meant a model-generated argument could switch off a safety behavior that only the operator should control.

The fix wasn't to patch each tool. I strip model-supplied suppression flags at one choke point: the client that every model-originated tool call goes through. I treat the model's arguments the way I'd treat user input from the internet. Whatever the model put in those flags, the client removes them before the call reaches real code, so turning off a fallback stays an operator decision.

Doing it in one place matters. If each tool had to defend itself, the next tool I add would be the one that forgets.

Contract three: what counts as an answer

This one came up before the migration, but the migration made it more pressing.

When the gateway between the bot and the model fails, it sometimes returns an error body as ordinary text: a status message, a short complaint about a quota or a timeout. Code that just reads "whatever text came back" will treat that as the model's output. I had a cycle that did exactly that. It took a gateway error as a conclusion, logged it like any other result, and moved on.

Now a reply has to look like an answer before it counts as one. If a cycle's output is empty, is shaped like an error, or is missing the structure the job expects, the cycle is marked degenerate. It exits with a failure status and pages me. A failed call is a failure, even when it arrives as a string.

The same idea showed up elsewhere that week: every way an autonomous cycle can end now has a typed outcome, and an outcome that wasn't recovered is never recorded as success.

Why "untrusted" is the right word

I used to think of the model as a capable colleague with access to the system. Now I think of it as an external service at the far end of an interface I control. Its inputs are something I build and account for. Its requests are untrusted parameters. Its responses have to be validated before anything acts on them.

None of this is specific to language models. It's how I'd treat any third-party API. The agentic tool hid that, because it made the model feel like it was inside the system instead of across a boundary from it. Taking that away made the boundary visible, and having to write the three contracts down is what made the change worth doing.

The old setup wasn't unsafe because the model was bad. It was hard to verify because I'd never said what the model was supposed to be doing, so I had nothing to check its work against.

Disclaimer: This journal documents a personal software-engineering project. The system described trades a paper (simulated) account. Nothing here is investment advice, a recommendation, or a signal, and no market data or trading performance is provided. Content is about building software.