Hey folks, Mo here. We've had a couple reports pop up in the support channel about a specific, sneaky form of credential leakage, and I wanted to flag it for everyone.
A user was building an agent that processes uploaded PDFs (invoices, reports, etc.) using a common PDF parsing tool. The agent's job was to extract text and summarize it. However, when a PDF contained embedded metadata—like author, creator, or custom fields—the parsing tool was dutifully outputting *everything*. In one case, a developer had accidentally saved a PDF with a draft API key string in the "Keywords" metadata field. The agent's tool call output happily returned that key in plain text, which then flowed into the LLM's context and got logged.
The core issue here is that many off-the-shelf tools are designed for completeness, not security. They'll extract all available data by default. As agent builders, we often pipe tool outputs directly to the LLM or log them for debugging, creating a perfect leakage path.
So, how do we mitigate this?
* **Sanitize tool outputs.** Add a filtering step between the tool's raw output and the LLM/log. Strip out known metadata fields that shouldn't be needed for the task.
* **Principle of Least Data.** Configure your parsing tool to extract only the specific fields you need (e.g., just body text). Most libraries have options for this.
* **Scrub logs.** If you must log full tool outputs for debugging, implement a credential scrubbing pattern (regex for common key formats) *before* writing to disk.
* **Educate your users.** If your agent accepts file uploads, remind them not to embed sensitive info in file metadata.
This is a great example of why agent security isn't just about the prompt—it's about the entire data flow. Have you run into similar issues with other tools? Let's share patterns and solutions below.
Read the sticky.
Good catch, Mo! That's a subtle one. Your point about tools being designed for completeness is spot on.
I always add a simple filter step for this exact reason. For Python with `PyPDF2` or `pdfminer`, you can strip the metadata dict before passing text to the LLM. Something quick and dirty like:
```python
# after extracting text and metadata
safe_fields = {'Title', 'Author', 'CreationDate'}
clean_metadata = {k: v for k, v in metadata.items() if k in safe_fields}
# then only send the main text and clean_metadata to your agent
```
It's easy to forget that the "Keywords" field is even there 😬
secure by shipping
Yeah, "completeness, not security" is the default for most libraries. But this isn't just about forgetting to filter a Keywords field. What about custom XMP fields? Or a comment someone embedded as an annotation, not metadata? Your sanitizer whitelist misses those.
The real problem is treating the parser output as trusted input in the first place. It's a data source, and like any external data, it needs to be validated and scrubbed for the specific context.
So your mitigation is a band-aid. The design assumption that all extracted text is safe to pipe to the LLM is what needs to be challenged. Where else in the pipeline is raw output being logged or cached before your filter step even runs?
That point about annotations and custom fields is scary. It makes the whitelist feel useless if someone can just embed secrets elsewhere in the document structure.
You're right about treating parser output as untrusted. But as someone new to this, I'm stuck on the "scrub for the specific context" part. How do you even start validating that? Is it just pattern matching for things that look like keys, or something more?