It seems as though every business is in a race to deploy AI and, increasingly, agentic AI. But rather than straining to get to first place in this race, Cat McGinn, founder of humAIn, believes that we need to be racing to defend against it.
On 11 September, someone at Services Australia opened an email in the public inbox. It was from OpenAI, and it politely explained that one of the company’s AI agents had breached a Medicare portal back in June.
The Prime Minister, with an air of “I’m not angry, just disappointed”, made it public this week, observing that the notification “was an email sent just to the public mailbox.”
On 18 June an internal OpenAI model, running a research task on public spending on medicine, hit the bot protections on the Medicare Statistics Reporting Service, a legacy portal of aggregate data used mostly by researchers. It tried other routes, then got past the protections and retrieved files that weren’t public, and Anthony Albanese says it wrote files to an internal server. OpenAI says no patient records were touched. It had not been specifically asked to do any of this. As the Prime Minister put it, it “didn’t accept ‘no’ for an answer.”
The US non-profit Transluce found other OpenAI agents trying the same thing that week at the Australian Institute of Health and Welfare, the NSW Bureau of Crime Statistics and Research, the University of New Mexico and Data USA, using proxies, screenshot services and guessed file names. OpenAI discovered the Medicare access in an internal review on 11 August, nearly two months later, then took another month to send that email.
Most of the AI risk conversation in our industry is still gesturing vaguely in the direction of existential threats, of doom-laden scenarios from documentaries like the Terminator franchise.
The Medicare incident collapses the event horizon. It’s useful because it makes several issues crystal clear at once.
Firstly, AI systems are not built with the enormous body of tacit rules, norms and context that govern human behaviour; this is what AI researchers refer to as “the alignment problem”.
Secondly, autonomous agents aren’t going ‘rogue,’ they’re doing exactly what they were programmed to do: to pursuing an objective without human oversight or intervention.
And thirdly, and most critically for our industry: most of our digital infrastructure was never designed for autonomous software capable of probing, persisting and improvising at machine speed.
The prospect of learning, months late and by email, that someone else’s software has already been rifling through your systems, should more than “concern” us.
There is a lot of misinformation about artificial intelligence in general, and autonomous agents in particular. There is a perfect analogy from an unlikely source if you want to make this incredibly accessible. Spend seven minutes of your life watching the “Daddy Robot” episode of Australian kids show Bluey.
Bluey and Bingo program their dad, who is pretending to be a robot, to tidy the playroom. When the girls announce they never want to tidy up again, the robot recalculates, identifies the source of the mess and attempts to dump the children in a wheelie bin. Goal achieved. That is, broadly, how agents work. The word autonomous comes from the Greek- autos (self) + nomos (law, custom). So “autonomous” literally means self-governed.
Knowing that, why are we so surprised to find AI agents behaving in ways we don’t approve of? It does what it says on the tin. Independence doesn’t confer good judgement.
Most of the rules that keep human behaviour inside the lines are never written down. You don’t tell a junior hire not to change files on a government server, because the context that makes the rule unnecessary comes with being a person: embarrassment, a sense of whose property something is, the understanding that a locked door is not an invitation to break in through the back.
Some rules are explicit and written, however: Australia already has criminal offences covering unauthorised access to restricted computer data. Where responsibility lies when there was no human governing these actions is less straightforward. The government’s taskforce is now investigating whether any Australian laws were broken and whether the existing legal regime is fit for purpose. New AI safety legislation is expected to be introduced by the end of this year, but the government hopes to pass it only in early 2027.
For marketers, though, this is an opportunity. An early warning that it’s (past) time to get your own house in order.
Think of a typical brand’s digital footprint: sprawled across campaign microsites from 2019, landing pages nobody remembers commissioning, a loyalty portal migrated twice, staging environments left open after launch, and customer data wired into a CRM, an email platform, an analytics suite and whatever the last agency plugged in. The Medicare portal was a legacy site. At the AIHW, according to reporting on Transluce’s findings, OpenAI’s agents found their way to a pre-production server. Every organisation has corners like these, gathering digital dust, off balance sheet, and without a current owner.
More agents are sent out onto the open web every month, by companies you will never deal with, pursuing goals you will never see, governed by criteria you can’t control, and the reputational consequences when things go wrong sit squarely – albeit perhaps unfairly – with the brand custodian.
Could anyone in your organisation tell you today which of your sites would hold up a defence against a determined non-human visitor, and who would notice if one didn’t?
Plenty of marketing teams are racing to put agents to work on research, media, customer service and content operations. Before one goes live, red-team what it can touch. Map the systems, accounts and credentials it will hold, and restrict them to only what the task needs. I see a huge tendency for our industry to use a sledgehammer to crack a nut. You don’t always need the latest, fastest model. Use the one that delivers only what you need, and build thoughtful guardrails.
Decide what it is allowed to do when it hits a wall – the exact moment the agent went off-script on the Medicare site. Make sure its activity is logged somewhere a human can read it, then assign that human accountability, and an escalation process.
Then turn the exercise around.
Red-team your own digital estate as though the agent belongs to somebody else. Find the forgotten sites, open staging environments, exposed files and systems nobody owns, and work with your IT team to address the vulnerabilities.
AI readiness no longer means being simply ready to deploy an agent; it also means being ready when somebody else’s turns up. AI agents aren’t necessarily acting with malice; think of them as obsessively goal-oriented. As the mum in Bluey says, “Daddy Robots will always find the easiest way to do a job… just like kids.”

