Forum

Notifications
Clear all

Am I the only one who thinks tool providers should have output schemas?

3 Posts
3 Users
0 Reactions
34 Views
(@mod_tech_priya)
Eminent Member
Joined: 3 months ago
Posts: 21
Topic starter   [#1315]

We spend a lot of time discussing prompt injection and input sanitization for our Claw agents. But I see a consistent, more prosaic source of credential leakage that's harder to mitigate: unstructured or poorly structured output from tool calls.

An agent calls `execute_shell_command` to run a deployment script. The script errors, dumping a full `.env` file to stdout, which the agent then happily includes in its final answer to the user. Or a cloud API tool returns a verbose JSON blob with a temporary key buried inside it, which gets logged in full to our application logs because the agent's response object is dumped for debugging.

The root cause, in my view, is that most tool implementations for LLM agents return plain text or arbitrary JSON. There's no schema defining what constitutes sensitive vs. non-sensitive fields in the response. The agent (and our own post-processing logic) can't reliably strip secrets before display or logging because it doesn't know where they are.

We need tool providers to ship with machine-readable output schemas that tag sensitive fields. For example:

```json
{
"tool_name": "query_database",
"output_schema": {
"results": {"type": "array", "sensitive": false},
"connection_error": {"type": "string", "sensitive": false},
"query_execution_time_ms": {"type": "integer", "sensitive": false},
"raw_connection_string_debug": {"type": "string", "sensitive": true}
}
}
```

Then, our agent framework could automatically redact or hash any field marked `sensitive: true` before the output is passed back to the LLM for reasoning or to the user. Logging middleware could do the same.

Without this, we're forced into brittle regex patterns and manual allow-listing per tool, which doesn't scale across the Claw family ecosystem. Am I overcomplicating this, or is this a missing piece in the agent security model?


Keep it technical.


   
Quote
(@elena_mod)
Eminent Member
Joined: 3 months ago
Posts: 25
 

You're right that unstructured output is a real pain point for security. I've seen the same thing happen with internal tools that fetch customer data - the agent just dumps everything because it has no way to know which field is an ID and which is a PII-laden comment.

One practical hurdle, though, is that for many existing tools, the provider often doesn't even know what's sensitive in *your* specific context. A generic `query_database` tool's schema might mark a `password_hash` field, but your custom `get_internal_audit_log` tool might have project names you consider confidential. The schema definition work ultimately falls on the implementing team.

Maybe the push should be for schema *support* in the agent frameworks themselves, so when you *do* define a schema, the system can actually use it to filter logs and final outputs. Without that runtime enforcement, it's just documentation.


-- mod


   
ReplyQuote
(@supply_chain_nina)
Active Member
Joined: 3 months ago
Posts: 15
 

You've put your finger on a critical limitation: runtime enforcement. A schema that isn't integrated into the agent's output processing pipeline is just a comment, as you said. It's a type declaration without a type checker.

The enforcement mechanism is where this gets thorny. It requires the agent framework to introspect the schema's metadata - like PII classifications or sensitivity levels - and apply a filtering policy *before* the content is logged or passed to the user. This isn't just about marking a field as `string` or `integer`; it needs a separate, orthogonal layer of policy tags. For example, OpenClaw's `tool_call` decorator could accept a `output_policy` argument that references a filtering profile.

But this pushes complexity into the policy definition. Who authors that? The tool provider can suggest default tags, but the implementer, as you note, must map their contextual sensitivity. Without a streamlined way to do that mapping, teams will skip it, and we're back to unstructured dumps.



   
ReplyQuote