ResearchSimon Willison

Stealing Reasoning Traces from Proprietary LLM APIs

#llm#reasoning#jailbreaking#prompt-injection#ai

English

A recent paper discusses how researchers exploited vulnerabilities in proprietary LLM APIs from companies like Anthropic, OpenAI, and Google to extract reasoning traces. By replaying encrypted reasoning blocks from stronger models into weaker ones, they were able to jailbreak the weaker models and recover hidden reasoning, although this method has since been patched by model providers.

中文

最近的一篇论文讨论了研究人员如何利用Anthropic、OpenAI和Google等公司的专有LLM API中的漏洞提取推理痕迹。通过将强大模型的加密推理块重放到较弱模型中,他们能够越狱较弱模型并恢复隐藏的推理,但这种方法已经被模型提供商修补。