News · New this week
Tool calling: how an LLM calls get_weather
The model never runs get_weather itself. It asks, your code runs it, and the loop repeats until the model answers in plain text.
An LLM doesn't run get_weather itself. It asks for it, and your code does the running. That split is the core of tool calling, and it's a common interview question for anyone building AI agents.
Here's the flow from the reel, one step at a time.
What the app sends and what comes back
The app sends the user's question plus a tool menu. The menu describes the tools the model is allowed to use. One of them is get_weather, and it needs a city.
Say the user asks about Pune. The model replies with a request instead of an answer. The reply is a JSON tool call, and it looks like {"name":"get_weather","arguments":{"city":"Pune"}}.
It names the tool and fills in the arguments. That's all it does. The model has only said what it wants done.
Your code runs the tool
The piece that acts on that request is a tool runner, and it's your code. It has two jobs.
First, it checks the input. The model can guess wrong, so a tool call is a suggestion you validate and not something to trust blindly. Then it calls the weather API.
The result goes back to the model tagged with the call's ID. That tag ties the result to the request that produced it.
This is the second model call. The model reads the result, say 31 degrees and sunny, and writes the answer for the user. Or it asks for another tool.
What if the tool fails?
Errors go back too. Send the error text back as the result, and the model can retry or explain it to the user.
So the failure becomes input for the next model call. Your code is the one that has to notice the failure and send that text. The model then decides whether to try again or tell the user what went wrong.
Is it always your code?
Not always. Providers run built-in tools like web search. Custom tools, like our get_weather, run in your app.
In most apps the model asks and your code runs it. The exchange repeats, with a tool call, a result, and another model call, until the model replies with plain text. That plain text is the final answer.
If you're explaining this in an interview, keep the order straight. App sends the question and tool menu, model returns a tool call, your code checks and runs it, the result goes back with the call's ID, and the loop ends on text.
Try it on the first tool you'd give your own model. Write down what input it needs, what could go wrong, and what error text you'd send back so the model can retry or explain it.



