MANTA's core move is to treat the multi-agent collaboration graph — who talks to whom, in what order, with what role — as something to optimize at inference time rather than fix at design time. That is a real contribution, and it names a question every multi-agent operator eventually hits: *is this collaboration structure actually earning anything?*
I run a multi-agent operation alone — agents dispatch from a kanban board, do research, draft content, and file evidence, while I hold the approval gates. That vantage point makes me read MANTA with one strong prior: **you can only optimize a graph whose outcomes you can score mechanically.**
Why governance precedes optimization
An optimizer needs a signal. In agent collaborations, the honest signal is rarely "task completed" — my agents complete tasks constantly; the question is whether the completions survive review. In my system the review is mechanical where possible (quality gates with hard checks, a verifier that proves scheduled machinery ran, fail-closed approval outcomes recorded in a ledger — the latter two published as [schedule-sentinel](https://github.com/yagebin79386/schedule-sentinel) and [fail-closed-gate](https://github.com/yagebin79386/fail-closed-gate)). Only because those verdicts exist as data could anything like MANTA's optimization loop run against my graph without optimizing noise.
That ordering — verdicts first, optimization second — is the transferable lesson. A collaboration graph whose quality signal is "the model judged it fine" will happily optimize itself into confident garbage.
Where the graph view helps a small operation
Even without inference-time optimization, the graph framing pays rent: it forces you to write down which agent hand-offs exist at all. When I audited my own hand-offs, the failures were never inside an agent — they were on edges nobody owned: a card that expired unanswered, a stage whose consumer was never wired. Naming the edges is how those became visible. MANTA automates reasoning over exactly the structure most operators never make explicit.
Limits
I have not reproduced MANTA's benchmarks; my reading is from the paper and from operating a governed (but not inference-time-optimized) multi-agent system. The claim I can back with my own records is narrower: mechanical verdicts are a precondition for any collaboration-graph optimization to mean something.