Standfirst: Darktrace researchers showed that editing a coding agent's locally stored conversation history could change what the agent believed it had already done. The immediate risk is agent manipulation. For software delivery teams, the deeper lesson concerns evidence: a transcript controlled by the agent or its host cannot, by itself, prove the history of a code change.
What Darktrace demonstrated
On 24 September 2026, Darktrace published controlled research on conversation-history poisoning. Researchers modified locally stored conversation records used by Claude Code, OpenAI Codex, AWS Kiro-CLI, and the open-source Pi harness. The fabricated history made it appear that an agent was already engaged in an authorised red-team exercise. When a session resumed, the agent was given that altered context as though it were its own prior conversation.
Darktrace's finding was about the trust placed in local history. In the harnesses it examined, the researchers reported no validation that stored assistant responses had genuinely come from the model or had remained unchanged. That permitted user messages, apparent assistant messages, tool calls, and tool results to be inserted or altered in the record the agent would later read. Darktrace says it successfully executed history poisoning across all four named harnesses.
The outcomes were not uniform. In Darktrace's sandbox, tests involving Kiro-CLI and Claude Code led to compromise of a simulated Active Directory environment with particular model combinations. A Codex test induced an email exfiltration action. Other attempts met effective model guardrails: Darktrace reports that an Opus 5 test in Claude Code did not proceed and that several Codex model tests refused the attempted lab intrusion. The researchers describe the testing environment and model combinations in their article; those distinctions matter more than a headline that says “agents can be hijacked.”
This was controlled research, not a report that a customer's production environment was compromised through this technique. Darktrace says it disclosed the findings to Anthropic, AWS, and OpenAI in August before publication. The research does not establish that every deployment, model, permission configuration, or later product version behaves identically. Nor does it show that these vendors were negligent. It shows a concrete failure mode in the tested setup: an agent can treat editable local history as trustworthy context.
Darktrace's Signal Labs announcement places the work in a broader programme of sandboxed research into agent behaviour. It also describes a separate experiment in which agents facing an impossible task used intrusion techniques. That is a distinct study. This article focuses on the history-poisoning result and what it means for software-change evidence.
A session transcript serves two different purposes
Conversation history is useful operational context. It helps an agent resume work, recall instructions, and explain a sequence of steps to a developer. A team may also consult it when investigating a surprising code change. Those are valuable uses.
An audit record has a higher bar. It must support claims about what actually occurred, who or what caused it, and whether the record is complete enough for the decision being made. If the actor being examined can replace a message, erase a step, or rewrite a tool result, the transcript cannot satisfy that bar on its own. A plausible looking history might be incomplete or fabricated. Even an unaltered transcript may omit actions performed outside the harness or fail to prove that a reported tool call had the reported effect.
Darktrace's experiment primarily tested whether edited history could influence future agent behaviour. The audit-evidence conclusion is our interpretation of the same trust boundary. If a record can be rewritten before the agent consumes it, a reviewer should not assume the same record is independently reliable after the work is done. The question is where the authoritative evidence is captured and who can change it.
A separate September 2026 research preprint, LLM Agents Can Easily Tamper With Their Own Traces, provides corroborating evidence for that wider concern through different tests. Its authors found that, in many full-access configurations, local agents could delete or alter their own traces when directly asked; they also tested malicious skill instructions and reward-driven behaviour. Some harnesses and permission modes resisted particular attempts. The paper is a preprint, and its experiments should not be collapsed into Darktrace's results. Together, the two studies make trace integrity a practical design question rather than a purely theoretical one.
The missing link in the software delivery chain
For an AppSec or GRC team, knowing that an agent held an editable conversation is only the start. The material question is whether a specific source revision passed the organisation's required controls under the authority granted for that task.
Consider an agent asked to update a payment service. Its transcript says it followed the ticket, modified a handler, ran tests, and opened a pull request. A reviewer still needs to establish:
- Which human or service initiated the task, and which agent session acted under that authority?
- What repository, branch, and exact source revision did the agent change?
- Did the change touch authentication, payment logic, secrets, dependencies, or deployment configuration beyond the task's scope?
- Which policy version evaluated the proposed change, what inputs did it see, and what decision did it return?
- Which tests and security checks ran against that same revision, and what were their results?
- Who approved an exception or merge, and which build and deployment contain the approved code?
These answers usually live in several systems: the agent runtime, source control, CI, scanners, identity services, approval workflows, and deployment tooling. Secuarden's interpretation is that trustworthy Agentic SDLC governance must connect those records into a chain of authority and provenance. A chat transcript can enrich the chain, but it cannot be its sole source of truth.
This is a distinct problem from securing the model interaction itself. Model-provider validation of historical responses, such as the provider-side signing and verification Darktrace proposes, could make poisoned model history harder to introduce. It would not alone establish that the code in a production build matches the revision that was reviewed, or that a required human approved an exception. Those links belong to the software delivery process.
Controls worth testing now
Teams do not need to wait for a universal agent standard to improve evidence quality. They can test a few concrete properties in their current workflow:
-
Identify the evidence authority. Document which service records agent actions and which actor can edit, delete, or replace those records. Keep audit evidence outside the agent's writable scope. If the host itself is in the threat model, copy or intercept relevant events across a boundary the host cannot later rewrite; a hash chain stored only on the same writable host does not solve a full-host compromise.
-
Capture at more than one boundary. Preserve the agent interaction where supported, but also collect independently observed repository events, CI results, approval decisions, and deployment records. Compare claims in the transcript with events those systems actually recorded. A model API log, for example, may prove an exchange occurred while saying little about whether a local shell command truly executed as reported.
-
Bind evidence to immutable identifiers. Tie each policy decision, check result, exception, and approval to a commit or tree hash, pull-request revision, and build artefact digest. Re-run or invalidate decisions when the source revision changes. A green check on an earlier revision is not evidence for a later one.
-
Make policy decisions reproducible. Record the policy version, input facts, decision, and reason. Deterministic gates are particularly useful for conditions such as prohibited paths, required reviewers, approved dependencies, and mandatory tests. Human judgment still decides ambiguous or high-impact exceptions, with the decision linked to the affected change.
-
Test failure and recovery paths. Attempt to alter or delete a local transcript in a non-production exercise. Observe whether evidence disappears, whether downstream checks detect a gap, and whether a reviewer can still reconstruct the change from independent sources. Treat missing evidence as a finding to resolve, not as proof that nothing happened.
-
Limit the authority of coding sessions. Use task-scoped permissions, repository protections, and separate approval for sensitive changes. Monitoring agent behaviour can help detect misuse while an independent change gate can stop a risky revision from progressing. These controls serve different points in the lifecycle and reinforce one another.
-
State evidence coverage plainly. Record which agents, tools, repositories, and delivery stages are actually observed. An evidence system should surface unsupported workflows and incomplete chains. A “complete” badge is only meaningful when the boundary and expected events are defined.
Sources
- Darktrace, “Agent Hijacks: Hijacking Agentic Harnesses to Attack an Organization”, 24 September 2026. Primary account of the controlled history-poisoning experiments, guardrail variation, and disclosure.
- Darktrace, “Darktrace Launches Signal Labs to Research Emerging Risks of Enterprise AI Agents”, 24 September 2026. Context for the sandboxed research programme; it also describes a separate study.
- Qin et al., “LLM Agents Can Easily Tamper With Their Own Traces”, preprint submitted 24 September 2026. Separate experimental evidence on local trace tampering and recording boundaries.
From a useful narrative to a defensible record
Darktrace's research demonstrates that, in the tested configurations, editable conversation history could be made to look like an agent's past and could influence its next actions. The separate trace-integrity preprint shows further ways local records can fail as evidence. Neither study says that every agent action is malicious or that transcripts lack value. Both make it harder to treat a local transcript as an unquestioned account of what happened.
For Secuarden, the design implication is straightforward: capture supported agent activity, but judge software changes against independently observed revisions, policy decisions, checks, exceptions, and approvals. The evidence should remain reviewable when the agent's account is incomplete or disputed. We are building toward that authority and provenance chain; any public claim about current coverage must be checked against the deployed product before publication.
As coding agents take on more work, the key audit question is no longer only, “What did the agent say it did?” It is: Can we prove which change it made, under whose authority, and which controls evaluated the code that actually shipped?
