An oracle, not a linter

When an agent writes the Go, "137 warnings, use your judgement" is useless as an acceptance signal. So I built GoRacle: a build gate that answers yes or no, and assumes the author will cheat.

A confession to start: some of the Go I ship these days is written by a model. Not most of it, and never unsupervised, but on the right kind of task, a well-scoped tool, a migration, a test harness, I will hand the keyboard to an agent, go do something else, and come back to a diff. The share is small and growing, and I think that is quietly true for most of the industry. Which is why the interesting problem is not whether the model can write the code. On the right task it can. The interesting problem is who gets to say yes.

When a person writes code, the acceptance signal is diffuse and mostly social: review, taste, the reviewer’s mood, how close it is to Friday. When an agent writes code, all of that collapses into whatever check you run at the end. And the standard tooling answer, a linter, hands back the one thing an autonomous loop cannot use: a hundred and thirty-seven findings across three severity levels, and an implicit “use your judgement.” The agent has no judgement. That is the arrangement. Judgement was supposed to be my half of the deal, and I am not reading a hundred and thirty-seven of anything at the end of every loop.

So at the start of August I built the tool I actually wanted, and because it refuses to negotiate I called it GoRacle. It is open source, Apache-2.0, and it has exactly one opinion, which is that it will not have opinions. Every check is binary. A finding either does not exist, or it fails the build.

A cartoon mascot for GoRacle: a muscular, angry blue bear in a ragged black hood and leather belt, holding a large spiked wooden club, drawn against a white background.
GORACLE The mascot sets the tone. There is no severity ladder on a spiked club.

A warning is a decision somebody already made

Here is the thing I had to admit to myself before any of the design made sense: I have never, not once, maintained a warning list that meant anything. Every project I have touched with a warnings budget had the same lifecycle. The list starts small and meaningful, somebody merges under deadline, the list grows, and within a quarter the warnings are wallpaper. A warning is a finding somebody decided not to act on. The decision already happened; the tooling is just archiving it.

Linters institutionalise that. Severity levels, suppression files, baselines that snapshot today’s mess so only tomorrow’s mess counts. Each mechanism is individually reasonable and collectively they add up to a gate you can talk your way through. And a gate whose verdict can be negotiated is not a gate, it is a conversation. Fine for humans, who at least feel shame. Useless for an agent, which will happily read “52 warnings” as a pass because nothing stopped it.

Two panels. Left, headed a linter reports: three horizontal bars labelled info 61, warning 52, error 24, captioned 137 findings, three severities, use your judgement. Right, headed an oracle answers: two buttons reading yes and no, captioned and when it is no: where, what, why, and the fix, nothing else.
THE SIGNAL The whole argument in one picture. A machine can act on the right side. Nobody can act on the left side, which is why nobody does.

GoRacle keeps the useful half of a linter, the detection, and replaces the reporting half with a verdict. When the answer is no, every finding carries four things in every output format: where, what, one paragraph of why the check reached its conclusion, and the fix. That is what an agent can act on, and honestly it is what I can act on at six in the evening too.

Steal the detection, keep the verdict

The part I want to be precise about, because it is the part I would side-eye in someone else’s announcement post: GoRacle detects almost nothing itself, on purpose.

Dead code is golang.org/x/tools’ reachability analysis, the same technique as cmd/deadcode. Unused declarations are staticcheck’s U1000. Duplication is dupl, in the fork the golangci team maintains. Documentation is revive. The security catalogue is gosec, the whole of it, one GoRacle rule per weakness. Known vulnerabilities are govulncheck in symbol mode, which only fires when your program can actually reach the affected function; that is the difference between an alarm and a to-do list. Mutation testing is gremlins. All of them run in process, as ordinary pinned dependencies over a single loading pass, so there is no binary on the PATH and no version skew between five tools installed five ways.

Writing my own detectors would have been the fun part, and it would have been wrong. These engines have years of edge cases in them, and they are maintained by the people who find those edge cases for a living. When staticcheck learns something new, my gate learns it with a version bump. What none of those projects wanted to build is the part I added: one configuration, one merged report, and a verdict that cannot be talked down. There is no key anywhere in .goracle.yaml that lowers the severity of anything. Configuration decides what is analysed. It never decides how loudly a violation counts.

The checks that assume the author is cheating

If the premise is that a model writes the code, then the checks themselves become a target, because the model is also very good at satisfying checks. Not maliciously. It just optimises whatever you measure, with a straight face. Three checks in GoRacle exist specifically because the naive version of them is gameable, and they are the reason I say oracle instead of linter.

Three columns under the heading three ways a green build can lie, and a check for each. First: the API drifted; rename a method, fix the one caller, everything still compiles, no test fires; caught by api/surface-changed, diffed against a reviewed baseline. Second: the tests cannot fail; the model wrote the code and its tests, they agree, that proves nothing; caught by mutation/survivor, mutate the code and rerun the suite. Third: the docs are hollow; a comment reading Foo does Foo is present, well formed, and completely empty; caught by contract/doc-mismatch, a blindfolded model tests the comment.
THREE LIES Three failure modes that all end in a green build, and the check that makes each one fail honestly instead.

The first lie is the drifting API. An agent that rewrites a function and its only caller can rename a method or drop a field, and everything still compiles, so nothing anywhere fires. Silent public-surface changes are the worst kind of defect because the build gets greener as the damage gets worse. So api/surface-changed pins the exported surface to a file with one line per symbol, recomputes the real surface on every run, and fails on any difference. Changing the API is allowed; it just has to arrive as a reviewed diff of that file, in the same commit.

The second lie is the test suite that cannot fail. When the same author writes the code and its tests, the oracle stops being independent: the two can agree while both being wrong, and a suite that passes no matter what the code does reads as proof while proving nothing. Mutation testing is the only answer I know of. gremlins changes the code one operator at a time and reruns the suite against each mutant, and a mutant the tests survive is a behaviour change nobody noticed, which fails the build. It re-tests the module once per mutant, so it lives in the deep profile and runs nightly rather than in anyone’s inner loop.

The third lie is my favourite, because the naive check practically begs for it. A documentation rule that wants a comment starting with the identifier’s name is satisfied, in one token, by // Foo does Foo. That is worse than no comment, because the absence would at least be visible. GoRacle’s answer is a check where a second model is shown only the comment and the signature, never the body, and writes a test from what the comment promises. If that test fails against the real implementation, the comment describes behaviour the code does not have, and the build fails. The comment has to be true, not merely present. Blindfolding the checking model is the entire trick; the moment it can see the body, it will helpfully agree with it.

Never in the binary, and you can count it

A quality tool that leaks into a production build is a dependency you did not order, so this was the first constraint and it is not a promise in a README. The gate runs from a _test.go file, from a separate command, or as a go.mod tool; none of those reaches go build output. The go command has no hook a library can install itself into, so the wrappers invert the order: goracle build runs the gate, and only if it passes does it hand over to go build. By the time the compiler starts, GoRacle has already exited. And one rule, gate/leaked-import, fails the build if any compiled file of the analysed module ever imports the gate, so the constraint is enforced by the tool it constrains.

A pipeline diagram: a box labelled the gate, every check in one pass, with an arrow labelled passes to a box labelled go build, only now, then an arrow to a box labelled the binary, zero goracle inside. A branch labelled fails drops from the gate to a box reading exit 1, the compiler never starts. Below, two shell commands: go list dash deps counts zero goracle packages in the binary, and twenty-two in the test closure.
THE ORDER The gate goes in front of the compiler, not inside it. The two commands at the bottom are the receipt: twenty-two gate packages in the test closure, zero in the binary.

Best of all, the claim is checkable in two shell commands. go list -deps . piped through a grep for goracle: zero. The same over the test closure: twenty-two. I trust a number I can recompute over any sentence I can write here, including that one.

Being wrong, honestly

A gate with no warning level has to be much more careful about being wrong, because every false positive is a stopped build rather than a line of noise. Most of the remaining design fell out of taking that seriously.

The taint-driven security rules (command injection, path traversal, SSRF) report what the analysis could not prove safe, which is not the same as unsafe. A program whose whole job is running a toolchain and opening the files it is handed will fire on its own architecture. GoRacle itself turns those three rules off, and says so in its own .goracle.yaml, in the repository, where you can read it. Some programs legitimately use MD5 for cache keys or must speak old TLS to a legacy device; the gate cannot read that intent and does not try. You annotate the line with a reason, or disable that one rule and write down why. I think of it as the difference between giving the gate a fact about your domain and negotiating with it. Facts are welcome. Negotiation is not in the protocol.

The fast mode makes the same kind of admission. It runs in about 170 milliseconds by skipping everything that needs the type-checked program, and instead of pretending, it prints each skipped check and the reason. A fast pass never means clean; it means clean as far as syntax can tell, and the report says exactly that. Even the exit codes refuse to blur: 1 means the code failed, 2 means the invocation failed, and a vulnerability scan that cannot reach its database is a 2, never a quiet pass.

Does it hold up? The gate has run on its own repository from the first week, everything on, and it has already embarrassed me once: the very first vulnerability scan flagged five reachable CVEs in the standard library and forced the toolchain pin up to a fixed release. Day one. On the tool whose job is to catch exactly that. I sat with it for a minute, and then I decided it was the best endorsement the design could get. The gate does not care that I wrote it, which was, after all, the entire specification.