The samples are small and the raw evidence is private, so you cannot reproduce these numbers from this repo. Treat them as a direction, not a law.
- Finding code: no token savings. Six code-location tasks used 4.9% more tokens with the map, for somewhat better answers.
- Planning: three big-picture questions, two runs each, graded blind. Answers with
brieffound 21 of 40 key facts against 12 without it, made 6 wrong claims against 11, cost 29% less and took 28% less time. Token counts were about the same. One caveat: the map carried earlier, checked findings that overlap the answer key. Carrying checked work forward is the point, but it is not independent discovery. - A held-out check (four tasks frozen before a change to the answers, three arms, one model, graded blind): source reading alone, and source plus the map before and after the change, found the same key facts (15.5, 16 and 16 points). So there was no gain. The map arms used 20% to 37% more tokens. They also made more claims the grader could not check without the map: its own status lines, and design rules that live only in the map. Source reading alone was cheapest for a known-file lookup. The consumers never used the new compact module card.
So don't expect a smaller token bill, or better answers on simple lookups. In the planning trial, the gain was fewer wrong turns.
Methodology notes: the grading was blind to the arm, as described above. The raw transcripts and answer keys are private and not published, so none of this can be re-run from this repository.