Using small models to break down problems
Decompose messy documents with a fast local worker model, verify in software, and escalate only the disagreements to a larger reasoner.
SectionsHand a large model an entire spec and it'll invent scope. Hand a small model one paragraph and an enum and it usually stays in its lane.
The four-step workflow
Here's the pattern I use on messy project documents:
- Parse in code. Headings, lists, table rows. Give every source unit a stable ID. If the document already has structure, don't chunk it by token count.
- Extract, don't summarize. A 3B-class instruct model pulls out atomic facts, requirements, and qualifiers (
must,maybe,phase 2). Run two independent passes. If they disagree, escalate. - Verify against the source. Another constrained call checks whether a qualifier got dropped, a field got invented, or the scope changed. Failures go to a larger model with *only* the disputed evidence.
- Ask a human last. Save it for when a guess would change the architecture, the cost, or the acceptance criteria.
What application code owns
Application code owns counts, IDs, retries, and “what stage are we in.” Models do narrow semantic work. That's the opposite of an autonomous agent chatting with itself until the context window fills up.
Related: how I think about eval and tracing once those workers are in production.