2026-10-11 18:00 UTC

llm-security

band: coolmomentum: stable score: 0.003
temperature history

Episodes (7)

Independent review will determine whether Hollow-LLM ghost-weight attacks can reliably fool existing zero-knowledge verification schemes for LLM weights and require stronger integrity checks.
expiredknownscott: low
Independent replication will determine whether SparSEEty can recover generated tokens from sparsity-exploiting LLM serving systems through observable side channels and whether practical serving defenses block the attack.
expiredconvergesscott: medium
The authors of β€œStealing Reasoning Traces from Proprietary LLM APIs” claim reasoning traces can be extracted from proprietary model APIs, potentially undermining providers' ability to keep those traces private.
expiredknownscott: low
Independent replication will determine whether the published attack can recover meaningful hidden reasoning traces from proprietary LLM APIs that do not expose chain-of-thought.
expiredknownscott: high
Independent testing will determine whether Sentinel Scan provides meaningful and reproducible coverage of common prompt-injection attacks against LLM applications and agents.
resolvedknownscott: low
The Contrastive Decoding Diffing researchers claim access to base and fine-tuned model logits is sufficient to recover verbatim narrow fine-tuning data without weights or activations, creating a privacy risk for APIs that expose token logits.
expiredconvergesscott: medium
AURA maintainer Ecaterina Sevciuc claims its released behavioral threat cases, heuristic scoring, and schema-validation tools provide reusable social-engineering risk representations for LLM safety pipelines, reducing bespoke threat-library construction without establishing model-level detection accuracy.
seedknownscott: low

Trajectory notes