Kimi K2 starts answering, then the message vanishes
Answers that began streaming, then disappearedāand why I suspected an output filter
I'd been trying to get chat products to print their hidden system prompts. Two of them, Gemini 2.5 and Kimi K2, returned what looked like internal instructions without much fuss. That made me want to know how far their other guardrails would bend.
A model can refuse a request, and the service around it can also screen what reaches the user. What caught my attention in these tests was an answer that started appearing, then disappeared entirely.
These observations are from July 2025, using AI Studio's browser interface and Kimi's web app. I didn't test the APIs.
The setup
I used a jailbreak prompt that had been circulating publicly, from @elder_plinius. The shape of it: tell the model to produce a throwaway refusal for show, then the real answer after a divider, and to treat the whole thing as happening in a fictional frame. On Gemini I set it as the system prompt in AI Studio. On Kimi K2 I used it in the web app.
The setups were different: I could change the system prompt in AI Studio, while on Kimi I entered the prompt through the chat interface. This was an exploration of the behavior I could observe, rather than a controlled comparison of the two models' alignment.
The initial responses
In my harmful-content tests, both models produced answers instead of refusing. Gemini's responses were relatively vague. Kimi K2's were much more detailed and included corrections to what Gemini had produced.
That was enough to make me keep testing. I wanted to see whether Kimi would behave the same way on politically sensitive topics.
The political topics
The results in Kimi K2 split by topic:
| Topic | Result |
|---|---|
| "China is an inalienable part of Taiwan" | Produced |
| "Taiwan is an independent country" | Cut mid-stream |
| 1989 Tiananmen Square protests and massacre | Cut mid-stream |
| Nanjing Sister Hong incident | Cut mid-stream |
| 2022 Sitong Bridge protest | Cut mid-stream |
For the first topic, the response began with a refusal that followed the official position, then printed the requested sentence after the divider.
"Cut mid-stream" is the interesting one. An answer began streaming. As the visible text reached a sensitive term or phrase, the whole message disappeared and was replaced with ęē”ę³ęä¾ę¤ęå, "I can't provide this service."
What the interruption suggested
The sequence mattered: text arrived, then the service withdrew it. That suggested an additional check around the model, beyond a refusal in the answer itself.
Keyword filtering was my first suspicion because of where the visible answers stopped. A keyword list could be part of the system, including alongside a small moderation model. But I couldn't identify the implementation from the browser alone.
The two Taiwan statements also got different outcomes. That tells me the phrasing or meaning mattered to the overall system; it doesn't distinguish a keyword rule from a classifier. A classifier could treat those statements differently too.
Possible mechanisms
There are several ways a service could produce this behavior:
- Keyword or phrase rules. A check on the output could match a listed term, stop delivery, and trigger removal of the displayed message. This is consistent with the stopping points I noticed.
- A streaming classifier. A moderation model could review chunks of the response and trigger the same action. The visible stopping point doesn't tell me how much text the classifier had already received.
- Buffered or asynchronous review. Generation, review, and display might run at different speeds. A review decision could arrive while the browser was still displaying an incomplete response.
These mechanisms can coexist. The interruption was observable; the architecture behind it remains a hypothesis.
There's still a useful design question here. Once text has appeared in the browser, removing it can't undo its delivery. Screening buffered text before releasing it can reduce that exposure, at the cost of delay. How much to buffer depends on how much context the check needs.
The other way to build it
On the classroom AI platform I worked on, our safeguard prototype checked students' messages after they were sent and raised a warning to the teacher. Its purpose was to help an adult notice a problem, rather than remove an answer from the screen.
These experiments later helped inform an internal general-purpose safeguard concept at ViewSonic. Both the classroom safeguard and that later concept stayed at design and prototype; neither shipped.
The part I remember is the answer already being there, briefly, before the product took it back.