For IT teams, the central issue is not whether an AI assistant can be made useful, but how to contain its authority once it is connected to enterprise systems. The article highlights a familiar platform-design problem: the more an assistant can read, write and act across email, files, calendars and web services, the more it behaves like a privileged integration layer rather than a simple chatbot. That changes the security model from โprotect the modelโ to โprotect every action path.โ
This makes deployment architecture the real control plane. Running an assistant on isolated hosts, separate cloud environments or tightly scoped service accounts can reduce blast radius, but only if access is segmented per workflow. In practice, teams will need explicit boundaries for identity, storage, browser automation, payments and local file access, plus logging that can explain which instruction caused which action. Without that separation, rollback and incident response become difficult once an agent starts chaining tools together.
Prompt injection also exposes a deeper integration risk: untrusted content is now execution-adjacent. Email bodies, web pages, documents and chat messages are no longer just inputs to be summarized; they may influence downstream actions. That means existing controls such as allowlists, content filtering, approval gates and least-privilege access need to be adapted for agent workflows. The most reliable pattern is likely to be task-specific constraints rather than a single general-purpose assistant with broad rights.
For decision-makers, the trade-off is operational as much as technical. Stronger guardrails improve safety, but can slow automation, increase false positives and raise support overhead. Teams evaluating secure AI assistants should therefore treat them like any other high-trust platform service: pilot with narrow scopes, define revocation paths, test failure modes and align security, data governance and end-user experience before expanding access.
AI agents are a risky business. Even when stuck inside the chatbox window, LLMs will make mistakes and behave badly. Once they have tools that they can use to interact with the outside world, such as web browsers and email addresses, the consequences of those mistakes become far more serious.
That might explain why the first breakthrough LLM personal assistant came not from one of the major AI labs, which have to worry about reputation and liability, but from an independent software engineer, Peter Steinberger. In November of 2025, Steinberger uploaded his tool, now called OpenClaw, to GitHub, and in late January the project went viral.
OpenClaw harnesses existing LLMs to let users create their own bespoke assistants. For some users, this means handing over reams of personal data, from years of emails to the contents of their hard drive. That has security experts thoroughly freaked out. The risks posed by OpenClaw are so extensive that it would probably take someone the better part of a week to read all Original Postersonal-ai-agents-like-openclaw-are-a-security-nightmare">of the security blog posts on it that have cropped up in the past few weeks. The Chinese government took the step of issuing a public warning about OpenClawโs security vulnerabilities.
In response to these concerns, Steinberger posted on X that nontechnical people should not use the software. (He did not respond to a request for comment for this article.) But thereโs a clear appetite for what OpenClaw is offering, and itโs not limited to people who can run their own software security audits. Any AI companies that hope to get in on the personal assistant business will need to figure out how to build a system that will keep usersโ data safe and secure. To do so, theyโll need to borrow approaches from the cutting edge of agent security research.
Risk management
OpenClaw is, in essence, a mecha suit for LLMs. Users can choose any LLM they like to act as the pilot; that LLM then gains access to improved memory capabilities and the ability to set itself tasks that it repeats on a regular cadence. Unlike the agentic offerings from the major AI companies, OpenClaw agents are meant to be on 24-7, and users can communicate with them using WhatsApp or other messaging apps. That means they can act like a superpowered personal assistant who wakes you each morning with a personalized to-do list, plans vacations while you work, and spins up new apps in its spare time.
But all that power has consequences. If you want your AI personal assistant to manage your inbox, then you need to give it access to your emailโand all the sensitive information contained there. If you want it to make purchases on your behalf, you need to give it your credit card info. And if you want it to do tasks on your computer, such as writing code, it needs some access to your local files.ย
There are a few ways this can go wrong. The first is that the AI assistant might make a mistake, as when a userโs Google Antigravity coding agent reportedly wiped his entire hard drive. The second is that someone might gain access to the agent using conventional hacking tools and use it to either extract sensitive data or run malicious code. In the weeks since OpenClaw went viral, security researchers have demonstrated numerous such vulnerabilities that put security-naรฏve users at risk.
Both of these dangers can be managed: Some users are choosing to run their OpenClaw agents on separate computers or in the cloud, which protects data on their hard drives from being erased, and other vulnerabilities could be fixed using tried-and-true security approaches.
But the experts I spoke to for this article were focused on a much more insidious security risk known as prompt injection. Prompt injection is effectively LLM hijacking: Simply by posting malicious text or images on a website that an LLM might peruse, or sending them to an inbox that an LLM reads, attackers can bend it to their will.
And if that LLM has access to any of its userโs private information, the consequences could be dire. โUsing something like OpenClaw is like giving your wallet to a stranger in the street,โ says Nicolas Papernot, a professor of electrical and computer engineering at the University of Toronto. Whether or not the major AI companies can feel comfortable offering personal assistants may come down to the quality of the defenses that they can muster against such attacks.
Itโs important to note here that prompt injection has not yet caused any catastrophes, or at least none that have been publicly reported. But now that there are likely hundreds of thousands of OpenClaw agents buzzing around the internet, prompt injection might start to look like a much more appealing strategy for cybercriminals. โTools like this are incentivizing malicious actors to attack a much broader population,โ Papernot says.ย
Building guardrails
The term โprompt injectionโ was coined by the popular LLM blogger Simon Willison in 2022, a couple of months before ChatGPT was released. Even back then, it was possible to discern that LLMs would introduce a completely new type of security vulnerability once they came into widespread use. LLMs canโt tell apart the instructions that they receive from users and the data that they use to carry out those instructions, such as emails and web search resultsโto an LLM, theyโre all just text. So if an attacker embeds a few sentences in an email and the LLM mistakes them for an instruction from its user, the attacker can get the LLM to do anything it wants.
Prompt injection is a tough problem, and it doesnโt seem to be going away anytime soon. โWe donโt really have a silver-bullet defense right now,โ says Dawn Song, a professor of computer science at UC Berkeley. But thereโs a robust academic community working on the problem, and theyโve come up with strategies that could eventually make AI personal assistants safe.
Technically speaking, it is possible to use OpenClaw today without risking prompt injection: Just donโt connect it to the internet. But restricting OpenClaw from reading your emails, managing your calendar, and doing online research defeats much of the purpose of using an AI assistant. The trick of protecting against prompt injection is to prevent the LLM from responding to hijacking attempts while still giving it room to do its job.
One strategy is to train the LLM to ignore prompt injections. A major part of the LLM development process, called post-training, involves taking a model that knows how to produce realistic text and turning it into a useful assistant by โrewardingโ it for answering questions appropriately and โpunishingโ it when it fails to do so. These rewards and punishments are metaphorical, but the LLM learns from them as an animal would. Using this process, itโs possible to train an LLM not to respond to specific examples of prompt injection.
But thereโs a balance: Train an LLM to reject injected commands too enthusiastically, and it might also start to reject legitimate requests from the user. And because thereโs a fundamental element of randomness in LLM behavior, even an LLM that has been very effectively trained to resist prompt injection will likely still slip up every once in a while.
Another approach involves halting the prompt injection attack before it ever reaches the LLM. Typically, this involves using a specialized detector LLM to determine whether or not the data being sent to the original LLM contains any prompt injections. In a recent study, however, even the best-performing detector completely failed to pick up on certain categories of prompt injection attack.
The third strategy is more complicated. Rather than controlling the inputs to an LLM by detecting whether or not they contain a prompt injection, the goal is to formulate a policy that guides the LLMโs outputsโi.e., its behaviorsโand prevents it from doing anything harmful. Some defenses in this vein are quite simple: If an LLM is allowed to email only a few pre-approved addresses, for example, then it definitely wonโt send its userโs credit card information to an attacker. But such a policy would prevent the LLM from completing many useful tasks, such as researching and reaching out to potential professional contacts on behalf of its user.
โThe challenge is how to accurately define those policies,โ says Neil Gong, a professor of electrical and computer engineering at Duke University. โItโs a trade-off between utility and security.โ
On a larger scale, the entire agentic world is wrestling with that trade-off: At what point will agents be secure enough to be useful? Experts disagree. Song, whose startup, Virtue AI, makes an agent security platform, says she thinks itโs possible to safely deploy an AI personal assistant now. But Gong says, โWeโre not there yet.โย
Even if AI agents canโt yet be entirely protected against prompt injection, there are certainly ways to mitigate the risks. And itโs possible that some of those techniques could be implemented in OpenClaw. Last week, at the inaugural ClawCon event in San Francisco, Steinberger announced that heโd brought a security person on board to work on the tool.
As of now, OpenClaw remains vulnerable, though that hasnโt dissuaded its multitude of enthusiastic users. George Pickett, a volunteer maintainer of the OpenGlaw GitHub repository and a fan of the tool, says heโs taken some security measures to keep himself safe while using it: He runs it in the cloud, so that he doesnโt have to worry about accidentally deleting his hard drive, and heโs put mechanisms in place to ensure that no one else can connect to his assistant.
But he hasnโt taken any specific actions to prevent prompt injection. Heโs aware of the risk but says he hasnโt yet seen any reports of it happening with OpenClaw. โMaybe my perspective is a stupid way to look at it, but itโs unlikely that Iโll be the first one to be hacked,โ he says.
Enjoyed this article? Sign up for our newsletter to receive regular insights and stay connected.

