Copilot prompt injection goes viral in your documents
All systems based on chatbots can be prompt injected. You can always tell the bot to do things it shouldn’t. The AI vendors try to put in guard rails. The guard rails sort of work for about two seconds. The original sin of chatbots is that data and instructions are mixed together in the same…
All systems based on chatbots can be prompt injected. You can always tell the bot to do things it shouldn’t. The AI vendors try to put in guard rails. The guard rails sort of work for about two seconds. The original sin of chatbots is that data and instructions are mixed together in the same stream. I’m sure you can all come up with an idea off the top of your heads on how to separate them. But the AI vendors have tried all the ideas and none of them work. Because data and instructions are mixed together in the same stream. Here’s a fun trick you can do right now with Copilot in an organisation, found by researcher Håkon Måløy in March: [blog post] Malicious instructions hidden in an externally shared document could make Copilot alter drafted or edited documents in Word and propagate the attack to new documents. You hide your prompt in a document, which can even be completely outside the target’s Copilot organisation. That infected document gets used as a source by Copilot for Word. If Copilot sees your prompt and uses it as instructions, it may manipulate the document the user is editing — change financial numbers, send company data out to the attacker, and so on. The new document may in turn become a carrier and spread the infection — even without the original infected document being present. You now have an AI worm. This works often enough that Måløy reported this to Microsoft as a security hole in March. He gave them a 90-day reporting window, Microsoft extended that twice to a total of 144 days, and they still didn’t have a solid fix. So Måløy posted about the hole on July 28th. In their press statement after Måløy went public, Microsoft tried to imply that they did have a fix: [Register] To address this class of risk, we use a defense-in-depth strategy with safeguards that block malicious instructions at multiple points and help keep tasks aligned with users’ requests. But Måløy says that no, Microsoft did not fix the hole: this scenario remains exploitable at publication. I have weighed that carefully. The coordination period agreed with Microsoft has been exhausted, and testing shows that no robust mitigation for the broader vulnerability class is currently available. Two mitigation attempts, including a model upgrade, did not close the class. I have therefore chosen to disclose at the class level rather than the payload level. Microsoft did patch bits of the problem — they worked around the original wording of Måløy’s prompt. Then Måløy changed the wording and got through again. Twice. Because prompt injection is fundamentally unfixable, by the nature of the chatbot. If you’re so profoundly foolish as to send all your documents through a chatbot that’s hooked up to do things, then prompt injections are going to get you. Stop trying to use chatbots for real work! Video — PodcastSource: Pivot to AI — Published — Category: Business