2026-10-11 17:09 UTC

model-security

band: coolmomentum: stable score: 0.032
temperature history

Episodes (5)

Independent replication will determine whether previous-token prediction can reconstruct hidden LLM prompts with near-exact fidelity and create a practical prompt-confidentiality risk.
expiredconvergesscott: medium
Independent testing will determine whether Steganeur's token-choice encoding provides a reliable and practically usable covert channel in natural-looking LLM-generated text.
expiredconvergesscott: medium
The paper’s author claims constraining fine-tuning to subspaces learned from trusted LoRA adapters can block malicious model updates while preserving useful adaptation, potentially adding a geometric defense against fine-tuning poisoning.
expiredconvergesscott: medium
The authors of β€œStealing Reasoning Traces from Proprietary LLM APIs” claim reasoning traces can be extracted from proprietary model APIs, potentially undermining providers' ability to keep those traces private.
expiredknownscott: low
OpenAI says it disrupted a coordinated model-distillation campaign attributed to Kimi, and Kimi's response plus any further disclosures or enforcement would establish organized cross-lab distillation theft as a recognized, actively policed frontier-model threat.
resolvedconvergesscott: high

Trajectory notes