Catalogue · living document

Tool-calling failure modes on local models — a field catalogue

Every way a local model has broken tool calling in one real agent, September 2026, with the mechanical guard that stopped each one. Observed on gpt-oss:20b, qwen3.5:9b and gemma4:12b through Ollama.

· 12 min

This is a catalogue, not an essay. Each entry is one way the local model driving my desktop assistant, Kora, failed to call a tool correctly — and the guard that was built after the failure was diagnosed from the traces. I keep it because the same failures keep coming back under new names, and because most of them are not fixable from the prompt: on a 12B to 20B model, an instruction that contradicts the model's habit does not hold. What holds is a mechanism.

Some context so the entries make sense. Kora runs on a single agent loop: one system prompt, about twenty tools (web search, file reads and writes in a sandbox, opening things on the PC, delegating to a cloud model, writing to a memory file), a per-tool confirmation policy, and a JSONL trace of every request, tool call, rejection and final answer. Every entry below was found by reading that trace, never by guessing.

Longer write-ups are linked where they exist. Dates are 2026.

How to read an entry

Symptom is what the user saw. Mechanism is what the trace and the code showed. Guard is what now prevents it — always something mechanical: a regex on the model's own output, a signature check, a budget, a field the app fills in itself. When a prompt change was tried first and failed, the entry says so.

1. The infrastructure was lying

No tool definitions were being sent at all — for two days

Symptom The model "opens Documents" fine and fails on "Images" four times in a row. Argument names guessed wrong, enums ignored, calls written in a made-up syntax.

Mechanism The streaming request path copied format, think and temperature to the IPC call by hand — and not tools. The model knew the tools only from prose in the system prompt and improvised the call format. The audit that morning had blamed the chat template and the model. The real prompt weighed 4,700 tokens; the same prompt rebuilt outside the app with tools weighed 10,400.

Guard One options builder shared by both request paths, and an end-to-end test that drives the streaming path with a stubbed transport and asserts the tools reach it. Rule learned: compare the size of the prompt the server actually received with what you think you sent, before diagnosing the model.

A timeout that cancels nothing

Symptom The user asks for one poem and receives two. VRAM stays high for a minute after the answer.

Mechanism Promise.race([loop(), timeout]) decides who answers the caller; it never aborts the loser. The loop kept generating for 43 seconds after the fallback model had already replied, streaming into a bubble nobody cleaned up.

Guard The timeout is an AbortController combined with the user's stop button, propagated down to the HTTP request. The round in flight is cut, not orphaned.

The last allowed round's tool call was thrown away

Symptom A task that needs eight rounds fails with an empty answer, although the eighth round contained the correct call.

Mechanism The round cap was checked before executing the round's tool calls, so a budget of 8 rounds was really 7 usable ones.

Guard Check the cap after executing the calls. The cap only forbids starting another LLM round.

A fresh round budget on every retry — 32 rounds for one message

Symptom Median 5 rounds per message, tail up to 37. The tail is where the cost and the VRAM pressure live.

Mechanism Four single-shot retry guards (see below) each called the loop again with a brand-new budget of 8.

Guard A total budget per user message; each relaunch gets what is left, and a relaunch with fewer than two rounds left is skipped entirely — it cannot succeed, it can only burn time.

2. The model does not respect the schema

Argument names drift, and drift again

Symptom write_file rejected 4 times with identical arguments: {path, text} instead of {filename, content}. Later file_path, file_name, name. For delegation: content, prompt, message, text instead of task.

Mechanism The same concept had two names across the tool set (path for reads, filename for writes). A model that gets the same rejection four times does not change its answer.

Guard An alias table applied before validation: an equivalent, unambiguous key is normalised, never rejected. A correctly-named key is never overwritten by its alias. Rejecting and hoping the model self-corrects is not a strategy — the trace shows it never did.

A cosmetic required argument blocks valid calls

Symptom open_site_search fails 6 times out of 7 in one exchange.

Mechanism The reason field only feeds the confirmation dialog; execution never reads it. It was required in validation, so a call that was otherwise perfect was thrown away.

Guard Required in the schema shown to the model (we still ask for it), optional in validation. A display-only field must not be able to veto an action.

The chat template drops required and enum

Symptom delegate_to_provider called with {provider: "chatgpt"} and nothing else. The model's own reasoning: "the signature is delegate_to_provider(provider, task?) — we must provide at least provider".

Mechanism Reproduced against Ollama with a minimal schema: gpt-oss's Harmony renderer prints every field as name: string, required or not, and the enum vanishes. The only channel that survives is the per-field description text.

Guard A registry test that fails if any truly required field lacks an "(OBLIGATOIRE)" marker in its description, or if any enum is not restated in words. Honest note: the marker is proven to reach the model, but the failure is intermittent and an A/B could not show the marker prevents it. Switching to a model whose template keeps the schema (gemma4) did.

Example values in descriptions get sent literally

Symptom execute_cleanup(plan_id: "p_3fa9c2d1") — the example from the description, not the real plan id returned thirty seconds earlier. Then "[Plan p_20240523_1030]". Then "current_plan_id_from_previous_step".

Mechanism Two things. A value in a description is a value the model will send. And tool results do not survive from one user turn to the next in this architecture, so at the "yes, go ahead" turn the id was simply not in the model's context — it could only invent one.

Guard No example values in descriptions. And never require at turn N+1 a datum that only existed in a tool result at turn N: the id became optional, resolved by the app to the last plan of this conversation. Safety never rested on the id — it rests on the confirmation dialog, which shows the real file list.

The memory tool is treated as a key-value store

Symptom Entries written to the memory file: no_vouvoyer, kairos_port, "Ne pas vouvoiement". The user had said "I don't want you to use vous with me".

Mechanism The model first sends {key: "vouvoyer", value: "no"}, gets rejected, then flattens its own pair into the entry field. Rewriting the description to forbid key/value pairs changed nothing.

Guard First a shape check (no spaces, snake_case, fewer than three words) that rejects once per turn with a concrete example of the expected form. Then the real fix: the text is no longer the model's to write. The app strips the trigger phrase from the user's own sentence and stores that verbatim. The model decides when and where; it no longer chooses the words.

3. The model says it did something it did not do

Narrated action, no call

Symptom "I'm launching the call now to create the file!" — end of turn. File untouched.

Mechanism A prompt rule against it existed and was extended twice. Reproduced a third time. A 20B model does not reliably obey a textual rule about its own output.

Guard A regex on the model's answer (future-tense action verbs), combined with a structural signal: zero tools used this turn. If both, one retry with "call the tool now". Never insist beyond one. Crucially this never reads the user's request to guess intent — that approach was tried in August and abandoned.

The tool call written as text

Symptom The final answer is literally {"name":"open_site_search","arguments":{"site":"images","query":"cats"}}. Nothing executed. Worse: the passive memory judge read that "answer" and recorded that the search had been executed.

Mechanism Structured tool_calls empty, JSON emitted in the content channel instead — a known template parsing failure.

Guard A narrow detector: the whole content parses as an object with at most two top-level keys, name plus arguments/parameters. Never a generic "looks like JSON" check — a legitimate answer may quote JSON. One retry.

A fake success with a fake proof

Symptom "Done, I moved those 12 files to the trash. [Tools actually called to produce this answer: Cleanup plan execution]" — zero executions in the trace, 14 files still on disk.

Mechanism The app appends a provenance note to assistant turns in the history, listing the tools really used. A marker placed in an assistant turn is a pattern the model learns to reproduce — here with a tool name that does not exist.

Guard The marker is never written by the model, so its presence in the model's output is a structural signal of fabrication: trace it, retry once, strip it regardless. Next step, not done yet: move the note to the following user turn, which is less imitable.

The delegated answer arrived, and the model said it had not

Symptom "Sorry, I hit a technical problem retrieving ChatGPT's answer" — 2,674 characters from ChatGPT sitting in the same turn.

Mechanism My own guard. Round 1's call was refused by a form check with a message starting "[Error: ...]"; round 2 succeeded; the model anchored on the first message. The passive memory judge then persisted a false "ChatGPT retrieval failed" anomaly. The guards had manufactured the exact class of problem they exist to prevent.

Guard Refusal messages that do not announce themselves as errors ("[TO FIX — the provider has NOT been called yet, nothing failed on its side]"), a failure detector that recognises that prefix, and a memory judge that no longer counts an attempt that was later corrected as a failure.

4. The model repeats itself

The same confirmed call, five times, after two refusals

Symptom open_path requested, approved, executed. Requested again, identical, approved again, executed again. Refused. Requested again. Refused again. Requested again. The user killed the app.

Mechanism The refusal message said "do not come back to this". Ignored twice in a row.

Guard Per-loop memory of every confirmed call's signature (name + sorted arguments). An identical call after a refusal returns the same refusal without a dialog; after a success, the same result without re-executing. Scoped to non-automatic tools — repeating a read costs nothing.

Duplicate side effects on "automatic" tools

Symptom "Play" reports success and nothing changes. Memory entries written twice, seconds apart.

Mechanism The signature guard above only covered tools that need confirmation. play_pause is a real toggle: two executions cancel out. write_memory persists, twice.

Guard An explicit opt-in list of automatic tools that still get deduplicated because they persist or toggle something. Not all automatic tools — repeating a search is harmless.

Seven paid cloud calls for one question, five byte-identical

Symptom "Give me feedback on the script I just shared" sent to ChatGPT, which sees no script and says so. Sent again, identical. Seven times. More spent in one turn than in the project's whole history with that provider.

Mechanism The delegation tool only forwarded the task string; the provider never saw the conversation. The description asked the model to "rephrase so it is understandable without the conversation". Measured across two models, ten runs each: the model copies the 1–3k characters of content into the task 1 time out of 10, and 0 out of 10. Four successive guards trying to force it all failed.

Guard The app attaches the context: the last exchanges of the thread (capped at 6,000 characters) and the user's current request travel with every delegation. The model only says what to do. Plus signature dedup on the delegation tool. Rule learned: if the model will not do something 9 times out of 10, stop writing guards and have the app do it.

5. The fallback costs money and cannot act

70% of fallback API calls traced to argument bugs

Symptom The cloud budget drains. The user believed there was a "complex task → Claude" rule. There is not: since the tool loop replaced JSON classification, the only triggers for the cloud fallback are an empty final answer or a timeout.

Mechanism 17 automatic escalations in the trace, 12 immediately preceded by a rejected tool call for a misnamed argument. And the fallback has no tools — it cannot perform the action the local model failed to perform. It can only explain, expensively, that it cannot.

Guard No escalation when zero tools succeeded and at least one call was rejected this turn; a fixed honest sentence instead, never model-generated. Escalation kept for the case it is actually good at: a purely textual task the local model could not answer.

The paid answer was thrown away, then paid for again

Symptom A delegation succeeds (290 characters from ChatGPT), the local model's final round comes back empty, the app escalates to Claude — paying twice and showing the user neither of the two answers it had.

Mechanism 6 paid delegations discarded in the trace history, 12,836 characters. The final round after several tool results is where a small model in no-thinking mode most often returns nothing.

Guard A third path before escalation, for "empty final with tools used": one local retry with no tools exposed (mechanically impossible to call one, and the prompt loses twenty definitions) with the tool results re-injected; if that fails too, show the paid result verbatim. Escalation only when no tool produced anything.

6. The tool misinforms the model about itself

The search excluded the current thread and never said so

Symptom Kora writes a text about autumn, then, asked to check past conversations, declares "I have never been asked to write about autumn."

Mechanism By design, the past-conversation search excludes the current thread (it is already in context). The empty-result message said "no past thread mentions X" — without mentioning the exclusion. No model can guess a scope you hide from it.

Guard The empty result states its own scope: "this search does NOT cover the current conversation — if the topic is in the messages above, reread them". This is a different family from every entry above: not a model error, an interface error.

The memory judge wrote diagnoses it had never seen

Symptom Five "technical anomaly" memory entries, four factually false: "ChatGPT call timed out" (it was an argument rejection; DeepSeek was never called), "cannot retry ChatGPT" (it had succeeded 21 seconds earlier).

Mechanism The after-the-fact memory judge received the user message and the final answer — nothing about which tools ran or failed. It converted the model's "I couldn't" prose into an invented mechanical cause. And the memory file is re-injected into every subsequent prompt, so a false "ChatGPT is broken" fed the next failure.

Guard The judge receives the turn's tool record (executed, outcome, rejected, reason) as the only source of technical truth, with the sentence that was missing: "your own answer above is not one". Mechanically, an anomaly entry is refused when no tool actually failed this turn.

7. Things that were mine

Three regex traps in one hour, two of them the same trap mirrored

Symptom "Look in ~/Kora and write me an inventory" routed to "pure conversation", tools withdrawn; the model honestly answers it has no file access. Then "/Kora" without tilde fails after "~/Kora" is fixed.

Mechanism JavaScript \b relies on \w, where accented letters do not count: no word boundary before the "é" of "écris-moi", so the keyword never matched. Then the end anchor I replaced it with failed in the opposite direction on "/Kora". Then the typographic apostrophe (U+2019) from the spell-checker matched nothing.

Guard Unicode lookarounds, apostrophe normalisation upstream of every pattern — and the thing that actually held: a structural rule that an explicit path or filename in the message forbids concluding "pure conversation" and implies read tools, whatever the keywords say. Keyword lists never converge; structural guards do.

What generalises

This page grows as the trace does. If one of these entries is yours too, I would like to hear how it broke on your side: contact.