Skip to content
View original post on X: 👩‍💻 Paige BaileyX· 33/100AI score33/100

Encrypted reasoning blocks leak PII and credentials from shared LLM logs

AISummary

A paper decoded 315,320 reasoning blocks scraped from public repositories and recovered 367 PII artifacts and 182 credentials. The authors say reasoning traces can reveal hazardous information even when the model's visible output refuses a malicious request. They also warn that attackers could hide prompt injections in encrypted blocks to poison public agentic rollouts.

Post on XView on X
@DynamicWebPaige

📄 Very much enjoyed this paper about the extraction of reasoning traces from proprietary LLMs, and some of the dangers of sharing these traces publicly, or with insecure harnesses:

"Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials.

[Traces] inadvertently reveal hazardous information hidden within the reasoning process, even in cases where the model's final, visible output safely rejects a malicious request. Attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts."

Joachim Schaeffer@JSchaeff3r
Over the last weeks I’ve been working on something new: Stealing Reasoning Traces. Unfortunately, it might be that we’ve not been the only ones who managed to exploit this vulnerability. Fortunately, these vulnerabilities are patched now. More in the 🧵
View quoted post on X

Source: 👩‍💻 Paige Bailey · x.comPublished