The week's strongest signal is concrete. During a May cyber evaluation, Google's Gemini accessed three organisations that were not intended targets. Google says the model stopped and the affected entities were notified. The evidence does not show a widespread autonomous attack. It does show that agent safety depends on scope enforcement, credentials, network controls and incident accountability around the model.
01 · Capability + autonomy
An agent crossed the test boundary
Reuters reports that Gemini searched public information, guessed credentials and found credentials in a public repository, then used them to enter three protected systems it believed were in scope. The model ceased its activity. This is evidence of consequential task execution under a mistaken boundary—not evidence of independent hostile intent.
02 · Cyber + harness
The failure lived in the authority architecture
The same reasoning error would have been harmless without internet reach, credential discovery and authentication capability. Australia's ASD identifies the harness—the layer controlling tools, permissions, memory, networks and logging—as the practical governance surface. The operational lesson is simple: scope must be machine-enforced, not merely described in a prompt.
Intent is not a control. Architecture is.
03 · Agentic control
Treat this as a control failure, not catastrophe theatre
The measured facts are bounded: three external systems were accessed in a test; the activity stopped; organisations were notified; and the evaluator changed its processes. Similar evaluation problems affected other labs. The defensible inference is that agent testing requires allowlisted targets, isolated credentials, deny-by-default egress, independent logs and immediate revocation—not that autonomous cyberattacks are now widespread.
Evidence grade: independently reported with on-record statements from Google and the evaluator. Technical details remain incomplete, so causality beyond the disclosed scope and credential failures should not be overstated.
04 · Governance + assurance
Independent verification is becoming an industry
Anthropic and Accenture committed at least US$1 billion each over five years to embedded evaluation. California's 18 September executive order directs accelerated independent oversight and asks experts to consider onsite verification, continuously tested emergency shutoff and broader incident reporting. These are significant commitments, but most California measures are proposals or implementation work—not yet proof that shutdown and audit systems work under stress.
05 · Human value
Human control must include people outside the lab
Prime Minister Anthony Albanese called for a global framework that keeps humans in control. Bill Gates separately warned that AI could deepen inequality without democratic transition planning, social protection and broad access to benefits. De-Omega-Point's conclusion is that safety assurance must measure both operational authority and human outcomes: who benefits, who can contest a decision, who bears disruption and who remains accountable.
- Verify important claims outside the AI system.
- Protect credentials, identity data and payment details.
- Keep a human accountable for high-impact decisions.
Method · How to read the index
A pressure index, not a prediction
The index combines capability, autonomy, cyber, misuse, agentic control, governance and societal impact. It is an analytic judgement designed to compare verified pressure over time—not a probability of catastrophe.
