One request, one reply, one confident guess — the model API request and reply, priced in tokens
A model API call is one request in and one reply out, priced and limited in tokens, and anything about your business that is not inside that request does not exist for the model — which does not stop it answering.
Scene 01
One request, one reply, one confident guess
- Watch
- Try it
- Predict
- Capture
What actually happens when a program "asks an AI" something? Less than you would guess: it opens an ordinary connection to a server, sends one block of text, and gets one block of text back. Watch that call get assembled for Shopfront — the same online store as the metrics and tracing curricula, seen this time from its support desk. First the request appears: everything the program chose to send, including the standing instructions it puts in front of every customer message, and the customer's own line. Then the reply comes back. Then the same call is measured — providers bill and limit text in small chunks of about four characters, called tokens, so the bar beneath the stage is this call redrawn with length standing for quantity. Then the price lands top-right. Finally the same call runs again with a harder question, and you should watch what the reply does about a record nobody put in the request.
Where this sits in Build a production AI agent (from one API call up)
Scene 01 of 2, in the One call act — Text in, text out. No memory. Check why it stopped.. A model API is a function from text to text: it knows only what is inside the request you send, it is billed by the token, and when it knows nothing it still answers fluently.
Up next. One call answered one question. The customer writes back — what does the model remember?
All 2 scenes in Build a production AI agent (from one API call up) · Every curriculum