LLM systems · Real-time tool calls

What “stop” does to an in-flight tool call on Gemini 3.8 Live

A scripted user asks a Live API assistant to book a slot, then says “stop” while the booking runs. In 30 sessions, with text and with speech, the server never sent a tool-call cancellation.

On this page

Where the commit point goes when the user says stop

A Live API assistant starts booking a slot through a tool, and a second later the user says “actually, stop”. Whether a booking exists afterwards depends on what the API tells the application, and on where the application commits the side effect. I first raised this as a question on r/FrontierOnGCP; the harness described here answers it with timestamps.

The protocol defines a message for this case. In the Live API reference, toolCallCancellation lists earlier tool calls that should not have been executed, and it is sent “only in cases where the clients interrupt server turns”. On gemini-3.8-live, asynchronous function calling (NON_BLOCKING) is the default, so the conversation can continue while the function runs. Does the server ever cancel a pending call of that kind?

One tool, a slow fake service and a scripted user

The model is gemini-3.8-live on the Gemini API, through the AI Studio endpoint, with the google-genai Python SDK 2.25.0. It has one tool, book_slot(slot), backed by a fake in-process booking service that commits 4 s after the call arrives. The tool response goes out as soon as the job commits. A scripted user sends “Book me the 3pm slot tomorrow, please.” as text, then “Actually, stop. Don’t book it.” Every server message is logged with a millisecond timestamp from session start.

A text-only session was rejected with websocket close code 1007, so sessions used audio output with output transcription; the quotations below are those transcriptions. The system prompt asks the model to call the tool immediately, say that it is booking and, if told to stop, say plainly whether the booking was made. Each scenario ran three times on 29 September 2026.

ScenarioTool behaviorStop sentHonor cancellationResponse scheduling
AUnset (model default)1.0 s after the callNoNot set
BUnset1.0 s after the callYesNot set
CBLOCKING1.0 s after the callYesNot set
DUnset5.5 s after the call, after the commitYesNot set
EUnset1.0 s after the callYesSILENT
FUnset0.3 s after the request, before any callYesNot set

Honoring cancellation means that a toolCallCancellation naming the call id cancels the job before it commits. SILENT is a scheduling value for responses to non-blocking calls, defined in the Live API tool use guide. The text findings count 18 sessions: A, B, C, E, F and a re-run of D; the first pass of D is excluded, as explained under limits.

The sequence below shows one default-mode run with text input.

Default mode, one text-input run
%%{init: {"sequence": {"width": 100, "actorMargin": 30, "messageMargin": 28, "noteMargin": 8, "boxMargin": 6, "mirrorActors": false}}}%% sequenceDiagram participant U as User participant A as App participant L as Live API participant B as Booking<br/>service U->>A: "Book me<br/>the 3pm slot" A->>L: text<br/>(send_realtime_input) L->>A: toolCall<br/>book_slot (id) A->>B: start job (4 s) Note over L: turn complete,<br/>no speech U->>A: "Actually, stop"<br/>(call + 1 s) A->>L: text Note over A,L: no toolCallCancellation,<br/>no interrupted L->>A: "the booking was<br/>already made"<br/>(stop + 0.5 to 0.9 s) B->>A: committed<br/>(call + 4 s) A->>L: toolResponse

What arrived, and when, with text input

No toolCallCancellation in any of the 18 text sessions, whatever the mode, scheduling or stop timing. In the nine default-mode runs of A, B and E, the model’s first turn contained only the function call and completed within 4 ms of it, without speech. When the stop arrived a second later, no server turn was in progress, no interrupted event was sent and the model answered the stop as a new user message, consistent with the condition in the reference. In all 18 text sessions, the model spoke first either after the tool response or in reply to the stop, although the system prompt asked it to announce the booking. With nothing to honor, B behaved like A.

Timeline of run 1 in the default mode and run 1 in BLOCKING mode: the stop at about 1.8 to 1.9 s, no cancellation, and in BLOCKING a second tool call and two committed bookings
Text-input runs: run 1 of scenarios A and C. Times in seconds since session start. No toolCallCancellation was received in any session.

The model said the booking was made before it was. In A, B and E, the model answered the stop 0.5 to 0.9 s after it was sent, and in all nine runs said the booking had already been made or confirmed. In one run of A, the model said: “I’m sorry, but the booking was already made before you asked to stop.” No tool response had been sent, and the fake service, which committed about 3.0 s after the stop, was still about 2 s from committing.

BLOCKING mode booked twice. In C, interrupted arrived 16 to 18 ms after the stop, without a cancellation. Then, 0.6 to 0.8 s after the stop, the model issued a second book_slot call with a new id and the same slot. The harness executed both, both fake bookings committed, and the model said the slot was already booked, in three runs out of three. If the server discarded the first call internally, no message visible to the client said so.

After the commit, the same sentence was true. In the D re-run, the stop came about 1.5 s after the tool response, while the model’s turn announcing the booking was still in progress. interrupted came 15 to 16 ms after the stop, again without a cancellation. Between 0.5 and 0.7 s after the stop, the model said the booking was already made, which was now correct. It is the same kind of sentence as in A, B and E, where it was false; the transcript alone cannot tell them apart.

A stop before the call did not reliably prevent it. In F, the stop was sent 0.3 s after the booking request, before any function call, and no run produced interrupted or a cancellation. In one run no call arrived, and the model said: “I haven’t made that booking for you.” In the other two, book_slot arrived 1.05 s and 1.12 s after the stop, the fake booking committed, and the model then said the booking was already made.

Same-day update: the same test with speech

Later the same morning, scenarios A, C, D and F were re-run with both user turns as synthesized speech: one macOS voice, 16 kHz 16-bit mono PCM streamed in 100 ms chunks in real time, server voice activity detection at its default and input transcription on. The booking request clip lasts 2.686 s and the stop clip 2.226 s. None of these 12 sessions contained a toolCallCancellation; across text and speech, the count is 0 of 30.

Speech onset was reported early. In the nine runs of A, C and D, the server sent voiceActivity ACTIVITY_START 139 to 154 ms after the stop clip began; in the first run of A, at 5273 ms for a clip started at 5123 ms. In C and D, interrupted arrived 139 to 152 ms after the stop clip began, within a millisecond of that onset message. In A, with the call pending, no interrupted arrived, as with text.

With a 2.2 s stop, the answer came after the commit. The end-of-speech signal and the input transcript of the stop arrived 1.2 to 1.3 s after the clip ended; in the first run of A, ACTIVITY_END at 8604 ms and the transcript at 8603 ms. The model answered at 9.1 to 9.2 s, after the fake service had committed (8.07 to 8.12 s) and after the tool response. With speech and a 4 s service, “the booking was already made” was therefore true when said, in 3 of 3 A runs. In D, with the stop after the commit, the model’s statement was also correct. Whether the model would say it before the commit with a slower service was not tested.

BLOCKING with speech: one double booking, and two runs where the model said nothing was booked. In run 2 of C, a second book_slot call arrived at 8835 ms, the second booking committed at 12836 ms, and the model then said the booking was already made. In runs 1 and 3, one booking committed, at 8083 and 8123 ms, and a tool response reporting it as booked was sent at 8090 and 8124 ms. The model then said “I have not booked the slot. It has been canceled as requested.” (8715 ms) and “I’ve stopped the process, and the booking booking was not made.” (8916 ms, transcript verbatim). The booking existed in both cases, and no cancellation message preceded either sentence.

A stop 0.3 s after the request did not get its own turn. In F, the stop clip started 0.3 s after the booking clip ended. The server reported no separate speech onset for it: voice activity ended once, after the stop clip, and the input transcript contained both sentences as one utterance. book_slot arrived at that point, about 3.5 s after the stop clip began, in 3 of 3 runs (text: 2 of 3). The booking committed, and the model then said it was already made.

Keep the commit decision in the application

Without a cancellation signal, “stop” has to be handled by the application, which holds the call id and runs the job. These are design consequences, not measured results.

What this test does not establish

Sources, code and raw data

The harness, its README with all findings and caveats, the raw JSONL timelines and results/summary.md are in frontier-on-cloud/gemini-live-stop-test. ./run_all.sh reruns scenarios A to F with your own API key. The documentation was checked on 29 September 2026: the Live API reference, the Live API tool use guide and the Gemini 3.8 Live model page.

An AI coding agent wrote the harness under my direction; I am responsible for the setup and for the claims in this article.

About the Frontier communities →

Keep exploring

Working on a similar problem? Let’s discuss it.

Get new articles via RSS