I've spent this year running the same kind of work over and over: sweep a parameter, quantize a tier, benchmark it against the field, decide what to ship. Signal, the AP quant lineup, the ROCmFP4 work, Halo Box. Different projects, same shape underneath, and the same problem every time: the runs live in whatever the agent happened to produce, in whatever directory it happened to write to, and six weeks later I can't tell you which config actually produced the number I quoted.
That's why I built Fieldwork Ledger, and it's now live in private beta.
The problem it's built for
Agents are good at running experiments and bad at being a system of record for them. Ask an agent to sweep a parameter overnight and you'll get results, in a dozen slightly different formats, scattered across sessions, with no shared sense of what counted as a fair comparison. The agent that ran Tuesday's sweep doesn't know what Monday's agent already ruled out.
Fieldwork Ledger gives that work one shared structure instead: a Product for the thing you're researching, Campaigns for the decisions you're actually trying to make, Experiments for the testable ideas under each one, and Runs for every individual attempt. Schemas are typed, with units and an optimization direction, and they can grow over time without breaking the runs already recorded against them. Two runs only get compared where they were measured alike: same schema, same basis for comparison, not just a number pulled out of context.
Built for agents, not just people
The whole point is that an agent does the work end to end: plans a run, executes it, records the result, while I set the direction rather than doing the recording myself. It's already wired for Claude Code, Codex, DeepSeek, Ollama, vLLM, llama.cpp, and custom harnesses like Marshall. Credentials for an agent are scoped and revocable, with expiration, and Fieldwork only ever sees what the agent explicitly records: prompts, weights, and code stay wherever they already live.
Above is roughly what a published result looks like to a reader: the headline number, what it's measured against, and the chart behind it, all typed and comparable rather than a screenshot pasted into a doc. That's the whole idea: an agent can run the sweep, but the record of what actually holds up should survive past the session that produced it.
It's in private beta right now, built first for the kind of work I already do (model evaluation, sweeps, ablations) but aimed just as much at industrial R&D and research labs running agent-assisted experiments that need to hold up later. If that's your kind of work, request beta access.