The manipulation half of the hypothesis is well-grounded. In controlled experiments on LLM-augmented security operations (the closest analog to a C2 decision-support loop), adversarially crafted log fields—user_agent, http_uri, payload, dns_query—acted as a functional injection substrate against GPT-4o-mini across 48 conditions (200 logs each; classification, summarization, remediation tasks) . Context Manipulation (S3) achieved 0.96 injection success for summarization under naive defenses and 0.38 even under constrained defenses; Persona Hijack (S2) suppressed classification at 68% naive and retained ~20% ISR under the strongest defense in remediation.
These findings generalize as a caution about any LLM placed in high-stakes operational loops: raw input channels must be treated as adversarial, provenance must be separated from instruction channels, and human review is recommended for high-stakes decisions . Frontier-model vulnerability is further corroborated: a specialized red-teaming agent achieved 60% ASR on GPT-5.5 and 26% on Claude Opus 4.8, with 3/14 and 10/14 models reaching 100% ASR on bio-risk benchmarks, and RL post-training doubling attack effectiveness (32%→60%, 12%→26%) .
No supplied evidence addresses stress or readiness monitoring. The argument for undetectability is mechanistically plausible: an LLM lacks physiological arousal, cortisol, or subjective value signals, so it would exhibit no stress biomarkers regardless of manipulation. But this is a category distinction (machine vs. human operator), not an empirical finding—no study measured detection rates of manipulated LLM advisories against any monitoring system. Undetectability also ignores other detection channels: behavioral drift, output distribution shift, provenance auditing, and cross-model consistency checks are all viable detectors that the evidence indirectly supports (defenses measurably reduce ISR, showing manipulation is at least partially detectable/preventable ).
Manipulation as attack surface: yes, with direct moderate-strength evidence. Undetectability by stress/readiness monitoring: plausible in principle (no physiological substrate exists in an LLM) but empirically untested and likely overstated, because manipulation is demonstrably detectable through output-level defenses. To falsify: replicate across multiple LLMs and show attacker-controlled inputs cannot alter advisory outputs; or demonstrate that stress-readiness metrics correlate with manipulated LLM decision quality (which the LLM's nature makes unlikely). What would change the conclusion: direct experiments pairing manipulated LLM advisories with real monitoring telemetry in operational C2 simulations.
Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.