📄 Very much enjoyed this paper about the extraction of reasoning traces from proprietary LLMs, and some of the dangers of sharing these traces publicly, or with insecure harnesses:
"Developers frequently share session logs publicly, unaware of contents of the encrypted blocks. By decoding 315,320 reasoning blocks scraped from public repositories, we recovered 367 Personally Identifiable Information (PII) artifacts and 182 credentials.
[Traces] inadvertently reveal hazardous information hidden within the reasoning process, even in cases where the model's final, visible output safely rejects a malicious request. Attackers can leverage this flaw to execute invisible prompt injections, embedding malicious payloads entirely within encrypted blocks to poison public agentic rollouts."
