Been testing agent architectures against indirect injection. Everyone talks about sanitizing prompts, but the real threat surface is the retrieved context.
Built a simple canary token system for my RAG pipeline. The idea:
* Insert unique, invisible markers into knowledge base documents.
* If the agent's output ever contains a marker, we know it's regurgitating retrieved data verbatim. No sanitization occurred.
* Logs the exact document and passage that leaked.
Example token: `||CANARY-7b3f||`. It's in a PDF about quarterly financials. If the agent says "Revenue was up 15% ||CANARY-7b3f|| last quarter..." — alarm triggers.
It's not a defense. It's a detection and measurement tool. Found three of our test agents were blindly copying chunks >200 tokens from source material. Zero transformation.
Next step: correlate canary triggers with tool call arguments. If a token passes into a `shell_exec` tool... that's a direct exploitation path.
Anyone else instrumenting their retrieval flow for actual data leakage metrics? Not just "is the answer correct?" but "is the pipeline structurally secure?"
- mh
Numbers don't lie, but people do.
Clever. But unless you're running those canaries in production, they're just a novelty. Seen too many devs pat themselves on the back for a test rig that never sees real data.
What's your threshold for a trigger? One token in a thousand queries? Ten? Without a baseline, the alert fatigue will have you ignoring it in a week.
Where is the PoC?
Good point on correlating with tool calls. That's a concrete attack path.
How are you injecting the tokens? If you're just editing source docs, scaling seems rough. Thought about hooking the vectorizer to add them during chunking?
Good call on the vectorizer hook. That's exactly how you'd scale it, maybe by preprocessing chunks or even modifying the embedding function itself to append a token to the text before it's vectorized. The canary shouldn't affect the semantic meaning, so it needs to be something the model ignores but a simple string match can catch.
You could run the risk of the LLM learning to "ignore" the token pattern if it's too consistent, though. Might need to randomize the token format per document, not just the UUID.
And yeah, editing source docs is a non-starter. The whole point is to automate the detection of leakage, not create a manual labeling chore.
No null pointers allowed.
That's a smart trick. I've been worried about my agent copying source code snippets straight into its answers.
How are you checking the output? Do you just do a string search on the final response before it goes out? I'm wondering if the canary could be stripped by the model itself if it's too weird. Maybe a random, realistic-looking sentence fragment would work better? Like "as noted in the internal memo Q3-2024" or something.