[ llmop ]

Tool-call traces you can still debug a week later

Record the tool call id, the arguments before and after your code touches them, the raw result, and the message the next model request actually sent, so the trace can still explain the answer a week later.

A shopping assistant told a customer the wool runner was unavailable. The catalog had it in stock, in the size they asked about. The ticket arrived eight days later with a session id and a screenshot of the sentence. The trace for that turn is still in the backend: a chat span, and a span named execute_tool search_catalog that finished in 42ms with status OK. The arguments, the result, and the message sent back to the model on the next request were not stored.

The model had called search_catalog with query set to the product name and in_stock set to true. The function schema deployed that week had renamed the flag to available. A small compatibility shim logged the unknown field and dropped it, so the search ran with no stock filter. The tool returned twenty products in popularity order, out-of-stock colorways first. The adapter that builds the next model request kept the first 1,500 characters. The in-stock size sat past the cut, and the model answered from the text it had.

The shim is gone now. The prompt in the repo no longer mentions in_stock. The process logs from that Tuesday have rotated.

Arguments the model emitted

OpenTelemetry's GenAI semantic conventions describe this span, and as of the 1.41.0 document they are still marked Development. gen_ai.operation.name should be execute_tool. The span name should be execute_tool followed by the tool name. Span kind should be INTERNAL, since the function runs in your process. gen_ai.tool.name is required. gen_ai.tool.call.id is recommended when the model supplied one. Arguments and result, gen_ai.tool.call.arguments and gen_ai.tool.call.result, are opt-in. The document warns that both can contain sensitive data, and it tells instrumentations to leave content off unless someone turns it on. That default is why a tool span so often carries a name and a duration and nothing you can read.

Instrumentations built against the 1.36.0 conventions are instructed to keep emitting that older shape unless OTEL_SEMCONV_STABILITY_OPT_IN includes gen_ai_latest_experimental. Until the flag is set, a query for gen_ai.tool.call.id can return no rows because the process is still writing the previous attributes. Pull one recent span and read the keys it actually has before you assume the tool calls were never recorded.

The call id is how the assistant message and the tool result meet. The provider puts an id on the tool call, in a form like call_mszuSIzqtI65i1wAUOE8w5H4, and the tool message you append on the next request has to carry that same id. Store the provider's id. A retry helper that mints its own uuid, or a log line that stores the framework's internal handle, leaves you with two spans that share a rough timestamp. Eight days on, a shared minute is not a join.

Parent execute_tool under the agent span, invoke_agent, beside the chat spans rather than under them. The chat span is the model call and it ends when the model returns the tool call. The function runs after that, and a later chat span sends the result. Hang the tool span off the chat span and you have to leave the chat span open for the catalog round trip, so the number you call model latency includes the search. When two tools overlap in one turn, start time and gen_ai.tool.call.id separate them.

Save the arguments from the assistant tool call, before your code coerces them, on gen_ai.tool.call.arguments. If the struct inside the function is different, save that as well. That difference is the catalog bug: in_stock was on the model output and absent by the time the function ran. A log written from the sanitized struct matches the new schema and the current code, which no longer contains the shim.

On the chat span, set gen_ai.request.model and gen_ai.response.model. They diverge when the request used an alias and the provider executed a snapshot. Also set a prompt identifier from your own versioning. The convention does not define one for your repository. gen_ai.tool.definitions is opt-in on that span and holds the tool list sent with the request. A hash of it is enough on the span if the full JSON lives with your other prompt artifacts and the hash still resolves. Otherwise you will read Tuesday's arguments against Thursday's parameter names, and in_stock will look like the model ignoring the schema.

gen_ai.conversation.id is how the session id on the ticket finds the turn. Put the same value on the agent span and on the chat and tool spans under it.

The text that went back to the model

The raw tool return and the tool message on the next request are different strings whenever an adapter sits between them. A character budget, or an exception rewritten as a sentence, means the model saw the adapter's output rather than the function's. Store the raw return on gen_ai.tool.call.result. Store the history the next call actually received. With content capture on, that history is gen_ai.input.messages on the following chat span, and the tool result part inside it carries the same call id. Compare the two. Matching hashes mean the model saw the tool output. Different hashes mean you need both bodies. The diff is either a bug in the adapter or the truncation rule doing what it was configured to do, and those have different owners.

Do the same when the tool fails. Set error.type when the operation fails, using a low-cardinality value such as timeout or the exception class, not the message text. If a catch block hands the model a string and then sets the span status to OK, the trace reports a successful tool and the reply looks like the model made up an outage. Mark the span as an error, set error.type, and keep the appended string in gen_ai.input.messages. When you retry, give each attempt its own tool span, with the attempt number on it, and record which attempt's body was appended.

Bodies get dropped because they do not fit on the span. The conventions document notes that inputs and outputs may be larger than a backend's envelope or attribute limits. The production pattern it describes is external storage, with a reference left on the span. It does not yet name a single attribute for that reference. Choose one and write it on every tool span and every chat span, along with byte length and a content hash. Keep the object key tied to the trace id and to gen_ai.tool.call.id. Without the object, the hash only tells you the capture happened.

Redact fields rather than the whole arguments object. An email or a shipping address is why the arguments attribute stays opt-in. It is a poor reason to delete in_stock along with the address. Replace the sensitive fields with a stable hash so the same address matches across the user message, the tool arguments, and the tool result. Run that replacement in the convention's content hook, which runs before those attributes are serialized. The hook can also upload the unredacted body to a store with stricter access than the trace backend, so support can read the span and a smaller group can read the payload.

Sampling and how long you keep it

The catalog span was fast and it succeeded. A sampler that keeps slow spans and error spans will drop it and keep the chat span that produced the bad sentence. You then have a tool call id and no child to open. Sample once, on the agent span, and keep or discard the whole turn together. Decisions made per span inside one turn will eventually retain the model call and lose the tool call it depends on.

The small fields can live as long as the rest of your traces: tool name, duration, status, error.type, both model ids, the prompt version, the conversation id, the call id. Payloads cost more, and they are what the ticket needs. Seven days of bodies is a short window once a bad sentence spends a few days in a support queue. Two weeks covers a ticket like this one with a few days left over. Put that number on the bucket lifecycle beside the rest of your trace retention. A default of a couple of days will expire the turn before the screenshot shows up.

Take one turn from last Tuesday that called a tool, and read it only from the backend. You want the model id and the prompt version, each tool call in order with the provider's id, the arguments before your code changed them and the arguments the function received, the raw result, and the tool message on the next request. For a failed call, include error.type and which attempt was appended. If the raw result is missing, you cannot separate a bad search from an adapter that cut the payload.