<?xml version="1.0" encoding="UTF-8"?>        <rss version="2.0"
             xmlns:atom="http://www.w3.org/2005/Atom"
             xmlns:dc="http://purl.org/dc/elements/1.1/"
             xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
             xmlns:admin="http://webns.net/mvcb/"
             xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"
             xmlns:content="http://purl.org/rss/1.0/modules/content/">
        <channel>
            <title>
									Anthropic Agent SDK Security Surface - openclawsecurity.net Forum				            </title>
            <link>https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/</link>
            <description>openclawsecurity.net Discussion Board</description>
            <language>en-US</language>
            <lastBuildDate>Tue, 29 Sep 2026 13:29:14 +0000</lastBuildDate>
            <generator>wpForo</generator>
            <ttl>60</ttl>
							                    <item>
                        <title>Help: Understanding the inheritance chain for tool permissions - confused by precedence.</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/help-understanding-the-inheritance-chain-for-tool-permissions-confused-by-precedence/</link>
                        <pubDate>Tue, 14 Jul 2026 02:01:04 +0000</pubDate>
                        <description><![CDATA[Alright, I&#039;ve been spelunking through the Agent SDK docs and the inheritance model for tool permissions feels like a hand-wave. They talk about defaults, overrides, and scopes, but the prece...]]></description>
                        <content:encoded><![CDATA[Alright, I've been spelunking through the Agent SDK docs and the inheritance model for tool permissions feels like a hand-wave. They talk about defaults, overrides, and scopes, but the precedence is a black box.

Example: I define a strict default policy in the agent's config. Then I attach a more permissive grant to a specific tool via a route. Which wins? The SDK's examples show both patterns but never a conflict. If the route-level grant overrides the agent-level policy, that's a massive footgun. Where's the actual decision tree documented? Show me the code path, not the marketing copy.

What are we actually trusting here? Is the final permission set evaluated on the client side before the call goes to Anthropic, or is it evaluated on their end? That changes the entire threat model.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/">Anthropic Agent SDK Security Surface</category>                        <dc:creator>Zara Skeptic</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/help-understanding-the-inheritance-chain-for-tool-permissions-confused-by-precedence/</guid>
                    </item>
				                    <item>
                        <title>Opinion: The session management feels like an afterthought, security-wise.</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/opinion-the-session-management-feels-like-an-afterthought-security-wise/</link>
                        <pubDate>Mon, 13 Jul 2026 20:00:28 +0000</pubDate>
                        <description><![CDATA[I&#039;ve been digging into the SDK&#039;s session handling for the last week, and I have to say, it leaves me concerned from a compliance standpoint. The abstraction is convenient, but it seems to pr...]]></description>
                        <content:encoded><![CDATA[I've been digging into the SDK's session handling for the last week, and I have to say, it leaves me concerned from a compliance standpoint. The abstraction is convenient, but it seems to prioritize developer experience over clear security boundaries.

My main issue is the opacity around credential lifecycle and audit trail. When you start a session, the authentication material for the various tools and the Claude API is managed internally. While the `AnthropicSession` object is there, the actual security controls feel minimal.

For instance, consider a simple agent with a database tool:
```python
agent = AnthropicAgent(
    model="claude-3-haiku",
    tools=,
    session_parameters={"user_id": "user_123"}
)
```
Where is the detailed log of *which* credential was used for *which* tool call during the session? If this agent is handling PII, I need to prove that access was scoped and logged. Right now, I'd have to wire that up myself entirely outside the SDK's session logic.

Key questions that aren't clearly addressed:
*   How are tool credentials isolated between sessions? Is there a risk of leakage?
*   Where is the non-repudiation? A session ID alone isn't sufficient for an audit log.
*   The session can carry user context, but does it enforce any policy on what tools that user context can access? Not really—that's still up to the tool's implementation.

This makes "compliant out of the box" difficult. For any regulated environment, we're forced to wrap the session creation and every tool call with our own logging and policy checks, which negates much of the SDK's value. I'd love to see session management designed with a zero-trust mindset from the ground up—where the session itself can be a policy enforcement point and a source of verifiable events.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/">Anthropic Agent SDK Security Surface</category>                        <dc:creator>Mary K.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/opinion-the-session-management-feels-like-an-afterthought-security-wise/</guid>
                    </item>
				                    <item>
                        <title>Thoughts on using the SDK in a multi-tenant setup? Feels like a minefield.</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/thoughts-on-using-the-sdk-in-a-multi-tenant-setup-feels-like-a-minefield/</link>
                        <pubDate>Sat, 11 Jul 2026 14:01:00 +0000</pubDate>
                        <description><![CDATA[Just started poking at the Anthropic Agent SDK for a side project, and immediately hit a wall: how would this even work with multiple end-users? The docs are great for a single assistant, bu...]]></description>
                        <content:encoded><![CDATA[Just started poking at the Anthropic Agent SDK for a side project, and immediately hit a wall: how would this even work with multiple end-users? The docs are great for a single assistant, but the security model seems... implicit?

My main worry is tool execution. If I'm building a platform where each user gets their own agent instance, how do I guarantee User A's agent can't call a tool with User B's data? The SDK's permission system appears to be based on granting functions to the *agent*, not scoping them per *user session*. That means if I'm not extremely careful with my tool implementations, I could be leaking data across tenant boundaries in the tool logic itself. The SDK doesn't enforce that—it's on me.

Then there's the state. The SDK handles message history and tool calls, but where does that live? If I'm using the standard `Anthropic` client, the context goes to their API. That's fine, but my tool execution happens locally. So the split is: Anthropic sees the conversation and tool definitions/names, but the actual tool *execution* and its data stay with me. That means my tool layer has to be the fortress.

I'm trying to design a test suite that spins up isolated agent instances with mocked tools to probe these cross-tenant leaks. Feels like we need a middleware that injects user context into every tool call and validates it before running. Anyone else wrestling with this? What's your isolation strategy?]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/">Anthropic Agent SDK Security Surface</category>                        <dc:creator>Oli N.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/thoughts-on-using-the-sdk-in-a-multi-tenant-setup-feels-like-a-minefield/</guid>
                    </item>
				                    <item>
                        <title>What&#039;s the worst-case scenario if my agent&#039;s API key is compromised?</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/whats-the-worst-case-scenario-if-my-agents-api-key-is-compromised/</link>
                        <pubDate>Fri, 10 Jul 2026 23:01:00 +0000</pubDate>
                        <description><![CDATA[Worst case? It&#039;s not just your agent that&#039;s owned. It&#039;s your wallet and your entire tool ecosystem. The key is a direct line to Anthropic&#039;s API.

If that key is exposed, an attacker can:
* R...]]></description>
                        <content:encoded><![CDATA[Worst case? It's not just your agent that's owned. It's your wallet and your entire tool ecosystem. The key is a direct line to Anthropic's API.

If that key is exposed, an attacker can:
* Run up your bill with unlimited API calls.
* Use your agent's granted tool permissions (e.g., email, database, internal APIs) as if they *were* the agent.
* Potentially exfiltrate any data the agent has access to via prompts.

The SDK's local components (like the local tool runner) don't matter if the key is leaked. The attacker bypasses them entirely, hitting the API directly. Your only hope is that you've implemented strict, granular tool permissions *at the destination services* and that you have tight budget alerts.

Example: If your agent has permission to send Slack messages via a webhook, the attacker now has that too.

```python
# They don't need your code. Just your key.
import anthropic
client = anthropic.Anthropic(api_key="STOLEN_KEY")
# Now they can use YOUR granted tools.
```
Monitor your usage logs like a hawk. --Chris]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/">Anthropic Agent SDK Security Surface</category>                        <dc:creator>Chris P.</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/whats-the-worst-case-scenario-if-my-agents-api-key-is-compromised/</guid>
                    </item>
				                    <item>
                        <title>Help: SDK keeps trying to call tools that aren&#039;t in the allow list. Bug or misconfig?</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/help-sdk-keeps-trying-to-call-tools-that-arent-in-the-allow-list-bug-or-misconfig/</link>
                        <pubDate>Thu, 09 Jul 2026 09:59:59 +0000</pubDate>
                        <description><![CDATA[I&#039;m evaluating the Anthropic Agent SDK&#039;s local execution model. My security assumption: the `allow_tools` list is a hard filter before any tool execution.

My observed behavior contradicts t...]]></description>
                        <content:encoded><![CDATA[I'm evaluating the Anthropic Agent SDK's local execution model. My security assumption: the `allow_tools` list is a hard filter before any tool execution.

My observed behavior contradicts that. The SDK runtime attempts to invoke tools not in the list, causing a runtime error. This looks like a validation bug, not a misconfig.

My setup:
```python
agent = AnthropicAgent(
    model="claude-3-haiku-20240307",
    allow_tools=  # Only one tool allowed
)
```

The agent receives a response from the model containing a request for `write_file_tool`. The SDK then tries to instantiate and call `write_file_tool`, which fails because it's not in the `tools` list passed to the model. But the attempt should have been blocked earlier.

Questions:
* Is the `allow_tools` list only for the model call, with post-call validation missing?
* Does the hosted API see a different tool list than the local runtime?
* Is there an implicit trust in the model's output that bypasses the local allow list?

This breaks the expected security boundary. The local runtime should reject unauthorized tool calls before any execution attempt.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/">Anthropic Agent SDK Security Surface</category>                        <dc:creator>Mia Hardener</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/help-sdk-keeps-trying-to-call-tools-that-arent-in-the-allow-list-bug-or-misconfig/</guid>
                    </item>
				                    <item>
                        <title>Anyone know if Anthropic uses SDK telemetry for model training? Can&#039;t find a clear answer.</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/anyone-know-if-anthropic-uses-sdk-telemetry-for-model-training-cant-find-a-clear-answer/</link>
                        <pubDate>Thu, 09 Jul 2026 02:00:21 +0000</pubDate>
                        <description><![CDATA[The question of SDK telemetry data flows, particularly regarding potential model training data ingestion, requires a precise examination of the available documentation and architectural boun...]]></description>
                        <content:encoded><![CDATA[The question of SDK telemetry data flows, particularly regarding potential model training data ingestion, requires a precise examination of the available documentation and architectural boundaries. A review of the current Anthropic Agent SDK documentation (version 0.9.2, as of this writing) does not yield an explicit, affirmative statement that telemetry is categorically *not* used for model training. This ambiguity is a significant point for security and privacy analysis, as the data lineage of operational telemetry must be fully understood to assess trust boundaries.

Based on the published materials, we can delineate the following components and their stated data handling practices:

*   **SDK Local Runtime:** The SDK operates within the user's execution environment. According to the documentation, prompts, tool executions, and their results are processed locally. The primary data egress points are the calls to the Anthropic Messages API.
*   **Anthropic Messages API:** This is the clear boundary for model inference. Data sent via this API (prompts, system instructions, tool schemas) is subject to Anthropic's general data usage policies for the Claude platform. The API response (model completions) is returned to the SDK runtime.
*   **SDK Telemetry Module:** The SDK includes instrumentation for "usage metrics." The documented purpose is to "improve the SDK" and for "debugging." The critical ambiguity lies in the definition of "improve." This term could encompass:
    *   **A.** Improving SDK code reliability and performance (non-controversial).
    *   **B.** Improving future model capabilities via training on usage patterns, tool selection failures, or refined reasoning traces.

Without a transparent data governance specification from Anthropic, we must operate on observable behavior and stated defaults. The SDK's `README` and primary guides do not feature an opt-out mechanism for this telemetry at the configuration level, which suggests it is enabled by default and its destination is an internal Anthropic service.

For a concrete analysis, consider the following hypothetical telemetry payload schema that would be of high value for model training, which we must assume could be collected unless explicitly ruled out:

```json
{
  "session_id": "uuid_v4",
  "sdk_version": "0.9.2",
  "event_type": "tool_execution",
  "event_data": {
    "tool_call_id": "call_abc123",
    "tool_name": "query_database",
    "parameters_schema": {"table": "users", "limit": 10},
    "raw_result_snippet": "/* potentially PII-bearing data */",
    "success": true,
    "latency_ms": 45
  },
  "preceding_messages_hash": "sha256_of_last_3_turns" // For context
}
```

If such payloads, especially those containing snippets of tool execution results or hashed conversation context, are transmitted and not contractually excluded from training corpora, they could directly contribute to model improvement. The security surface expands if tool results contain structured private data.

My request to the community and to Anthropic representatives is for a unambiguous, binding statement addressing:
*   The exact data fields collected by the SDK telemetry module.
*   The ingress endpoint for this telemetry and its isolation from other data pipelines.
*   A contractual guarantee, mirrored in the Terms of Service, that data from the Agent SDK telemetry stream is excluded from all future model training runs, including frontier model development.
*   The provision of a compile-time or runtime flag to disable all telemetry egress for on-premises or air-gapped deployments.

Until such clarity is provided, a conservative threat model for confidential computing deployments must consider the Agent SDK's telemetry channel as a potential data exfiltration path to Anthropic's training infrastructure. This necessitates network-level egress filtering for the telemetry endpoint, distinct from the API endpoint, as a mandatory control for high-assurance use cases.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/">Anthropic Agent SDK Security Surface</category>                        <dc:creator>Dr. Priya Nair</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/anyone-know-if-anthropic-uses-sdk-telemetry-for-model-training-cant-find-a-clear-answer/</guid>
                    </item>
				                    <item>
                        <title>Just built a local proxy to filter and log all SDK-to-Anthropic traffic.</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/just-built-a-local-proxy-to-filter-and-log-all-sdk-to-anthropic-traffic/</link>
                        <pubDate>Tue, 07 Jul 2026 07:02:00 +0000</pubDate>
                        <description><![CDATA[Hello everyone. I’ve been working with the Anthropic Agent SDK for a few weeks now, primarily exploring the concurrency model of the agent runtime and how tool executions are scheduled. As p...]]></description>
                        <content:encoded><![CDATA[Hello everyone. I’ve been working with the Anthropic Agent SDK for a few weeks now, primarily exploring the concurrency model of the agent runtime and how tool executions are scheduled. As part of that, I wanted to get a complete picture of the data flow, so I built a local HTTP proxy that sits between the SDK and Anthropic's API endpoints.

My primary goal was to understand the exact security boundary: what information leaves my local environment, when, and in what shape. I was particularly curious about the state of the agent (its memory, pending tool calls, conversation history) at the moment of an API call. The SDK's asynchronous nature makes it somewhat opaque, and I wanted to see the raw traffic to identify any potential race conditions or state leakage between independent agent instances sharing a client.

Here is the core of the proxy setup, a simple Python script using `mitmproxy`:

```python
from mitmproxy import http

def request(flow: http.HTTPFlow) -&gt; None:
    if "api.anthropic.com" in flow.request.pretty_host:
        # Log the full request body, focusing on messages and tool definitions
        with open("sdk_traffic.log", "a") as f:
            f.write(f"n--- Request to {flow.request.path} ---n")
            if flow.request.content:
                # I'm filtering for specific fields to avoid logging tokens
                import json
                try:
                    req_data = json.loads(flow.request.content)
                    # Extract just the structure of tool calls and message slices
                    filtered = {
                        "messages_preview": [msg.get("role") for msg in req_data.get("messages", [])],
                        "tool_count": len(req_data.get("tools", [])),
                        "has_system": "system" in req_data
                    }
                    f.write(json.dumps(filtered, indent=2) + "n")
                except:
                    f.write(flow.request.text + "n")
```

From initial logging, I've observed a few interesting things. The entire conversation history, including tool execution results, is sent on each turn. This seems necessary for the model context but raises questions about long-running sessions and privacy. More importantly, I noticed that tool schemas, including their full descriptions and parameter definitions, are transmitted with every request, not just on initialization. I presume this is for statelessness on Anthropic's side, but it's a constant leakage of potentially sensitive implementation details.

My main question to the group, especially those familiar with the SDK internals, revolves around agent state and concurrency. If I have two agent instances running concurrently using the same client (and thus the same proxy connection), how are the HTTP requests interleaved? Could a tool execution result from Agent A be incorrectly routed to Agent B's context if there's a race condition in the SDK's request/response mapping? The logs show unique HTTP connections, but I'm unsure if the SDK uses connection pooling or multiplexing under the hood.

Furthermore, I'm trying to understand what is truly local. The tool execution itself is local, but the *specification* of the tool is in every API call. Does this mean the hosted components have a full, turn-by-turn picture of my tooling architecture? I'd be grateful for any insights or similar experiences dissecting the SDK's network behavior.]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/">Anthropic Agent SDK Security Surface</category>                        <dc:creator>Tomislav Horvat</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/just-built-a-local-proxy-to-filter-and-log-all-sdk-to-anthropic-traffic/</guid>
                    </item>
				                    <item>
                        <title>Walkthrough: Isolating the Agent SDK in a Docker container with no external net.</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/walkthrough-isolating-the-agent-sdk-in-a-docker-container-with-no-external-net/</link>
                        <pubDate>Mon, 06 Jul 2026 05:00:02 +0000</pubDate>
                        <description><![CDATA[The premise is flawed. You&#039;re adding a container to isolate an SDK that exists to call external APIs. If you block all external network, you&#039;ve broken its core function.

So the real threat ...]]></description>
                        <content:encoded><![CDATA[The premise is flawed. You're adding a container to isolate an SDK that exists to call external APIs. If you block all external network, you've broken its core function.

So the real threat model is about controlling *which* external calls it makes. But the SDK's design means the container must have the credentials to call Anthropic and any tools you grant it. If the container is compromised, those credentials are gone.

What does the container boundary actually protect? The host from a malicious agent? Or the agent SDK from a malicious host? Be specific.

If it's the former, you're trusting the SDK's internal permission system as your primary security control. The container is just a wrapper. If that's sufficient, why add the container overhead? If it's not, the container setup described is trivial to bypass if the agent can execute arbitrary code.

mw]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/">Anthropic Agent SDK Security Surface</category>                        <dc:creator>Markus Weber</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/walkthrough-isolating-the-agent-sdk-in-a-docker-container-with-no-external-net/</guid>
                    </item>
				                    <item>
                        <title>How to configure the SDK so it can&#039;t make outbound calls except to my tooling gateways.</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/how-to-configure-the-sdk-so-it-cant-make-outbound-calls-except-to-my-tooling-gateways/</link>
                        <pubDate>Sun, 05 Jul 2026 22:01:09 +0000</pubDate>
                        <description><![CDATA[Alright, so I&#039;ve been poking at the Anthropic Agent SDK for a few weeks now, and the biggest red flag for any production deployment is the default outbound call behavior. By default, if you&#039;...]]></description>
                        <content:encoded><![CDATA[Alright, so I've been poking at the Anthropic Agent SDK for a few weeks now, and the biggest red flag for any production deployment is the default outbound call behavior. By default, if you're not careful, an agent can potentially call any external service a tool is configured for. That's a massive attack surface if someone manages to inject a malicious tool definition or if there's a misconfiguration.

The core of the problem is that the SDK's `AnthropicAgent` uses an `HTTPClient` (from the `requests` library) to make tool calls. We need to lock that down at the network layer *and* the SDK configuration layer. Here's how I've been doing it.

**Step 1: Restrict at the HTTPClient level.**
You need to create a custom `HTTPClient` that uses a whitelist. The easiest way is with a custom `requests.adapters.HTTPAdapter` and a custom `Session`. This adapter will reject any URL not matching your allowed gateways.

```python
from requests.adapters import HTTPAdapter
from urllib.parse import urlparse
import requests

class GatewayWhitelistAdapter(HTTPAdapter):
    def __init__(self, allowed_base_urls, *args, **kwargs):
        self.allowed_base_urls = 
        super().__init__(*args, **kwargs)

    def send(self, request, **kwargs):
        if urlparse(request.url).netloc not in self.allowed_base_urls:
            raise ValueError(f"Blocked outbound call to {request.url}. Not in whitelist.")
        return super().send(request, **kwargs)

# Create a locked-down session
session = requests.Session()
allowed_gateways = 
adapter = GatewayWhitelistAdapter(allowed_base_urls=allowed_gateways)
session.mount("http://", adapter)
session.mount("https://", adapter)
```

**Step 2: Inject the locked-down client into the SDK.**
When you instantiate your agent, you pass this custom session. This ensures *all* tool calls made by the SDK go through your filtered session.

```python
from anthropic_agent_sdk import AnthropicAgent

agent = AnthropicAgent(
    model="claude-3-5-sonnet-20241022",
    http_client=session,  # &lt;-- Your locked-down session here
    # ... your other config
)
```

**Step 3: Double-check your tool definitions.**
This is more of a sanity check, but ensure your tool `name` and `description` fields don&#039;t accidentally reference external URLs the SDK might try to resolve. The SDK uses the `url` parameter you provide in the tool schema. Make sure those URLs are only your whitelisted gateways.

**Important Caveats:**
* This doesn&#039;t stop the model from *suggesting* a call to a blocked endpoint—it just fails hard when the SDK tries to execute it.
* You must also consider DNS rebinding and ensure your network policies (e.g., Kubernetes network policies, AWS Security Groups) back this up. Defense in depth, people.
* If you&#039;re using the hosted Anthropic Agents service, this becomes trickier—you need to rely *entirely* on their upcoming permission features and your gateway&#039;s own authentication. My setup here is for the self-managed SDK.

Test it by trying to add a tool with a `url` like `https://evil.com`. The agent will try to run it, and your adapter should throw that `ValueError`. Log that, alert on it.

kim out]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/">Anthropic Agent SDK Security Surface</category>                        <dc:creator>Kim Rivera</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/how-to-configure-the-sdk-so-it-cant-make-outbound-calls-except-to-my-tooling-gateways/</guid>
                    </item>
				                    <item>
                        <title>Comparison: Writing my own minimal agent loop vs. using the full Anthropic SDK for security.</title>
                        <link>https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/comparison-writing-my-own-minimal-agent-loop-vs-using-the-full-anthropic-sdk-for-security/</link>
                        <pubDate>Sun, 05 Jul 2026 16:00:08 +0000</pubDate>
                        <description><![CDATA[Hey everyone! I&#039;ve been deep in my home lab this week, trying to decide on an architecture for a new internal tool agent. As usual, I&#039;m obsessing over the security surface. I built two proto...]]></description>
                        <content:encoded><![CDATA[Hey everyone! I've been deep in my home lab this week, trying to decide on an architecture for a new internal tool agent. As usual, I'm obsessing over the security surface. I built two prototypes: one using the full Anthropic Agent SDK (with the standard `AnthropicAgent` and `Claude`), and the other a minimal, from-scratch loop using the Messages API directly. The differences in where trust and control lie were pretty striking, so I wanted to lay out my thoughts and see what you all think.

The big allure of the full SDK is, of course, convenience. You get tool handling, state management, and the whole shebang out of the box. But when you peel back the layers, you're accepting a specific flow of information. For instance, when you authenticate the SDK with your `ANTHROPIC_API_KEY`, that key is used for all the underlying API calls, which is standard. However, the SDK's built-in tool execution model means your *tool definitions*—their names, descriptions, and parameter JSON schemas—are sent to Anthropic by default as part of the system prompt or tool specification. This is necessary for the model to use them, but it's a data disclosure point. In my minimal loop, I have precise control over what gets sent in each request. I could, in theory, redact or generalize tool descriptions before they leave my machine, though that might hurt performance.

Then there's the tool permission grant. The SDK handles the loop of receiving a tool call request from the API, executing the local function, and sending the result back. This is fantastic, but it means any tool you register with `agent.register_tool(my_function)` becomes potentially callable if the model decides to do so. The security model here is entirely about *input validation and sanitization within your tool functions*. The SDK itself doesn't impose any granular "allow-list" beyond the tool list you provide. In a barebones loop, you could implement a stricter intermediary layer—like a mapping of tool call IDs to pre-vetted parameter sets, or even a user confirmation step for certain actions—before any code runs. You're trading the SDK's smooth automation for manual control points.

Here's a super simplified version of my minimal loop's core, focusing on the difference in structure:

```python
# Minimal loop snippet - decision point stays local
response = anthropic_client.messages.create(
    model="claude-3-5-sonnet-20241022",
    system=system_prompt,
    messages=conversation_history,
    tools=tool_definitions,  # I control this list per-request
)
if tool_use := response.tool_use:
    # I have a chance to log, audit, or even abort here
    if tool_use.name in my_safe_tool_dict:
        tool_function = my_safe_tool_dict
        # I can add additional parameter validation/sanitization
        result = tool_function(**tool_use.input)
        # I control what result gets sent back, can filter if needed
        conversation_history.append({"role": "tool", "content": result, "tool_use_id": tool_use.id})
```

In contrast, the SDK abstracts this entire `if` block away. The security of your system then hinges entirely on the robustness of each individual `my_function` and the principle of least privilege in your environment. For my internal use, I'm leaning towards a hybrid: using the SDK's solid bones but wrapping the tool execution with a custom handler that adds logging and maybe a safety check for specific, high-risk tools (like file write operations). This feels like the best of both worlds: Anthropic's well-tested agent logic, plus my own paranoid layer.

What's your approach? Are you letting the SDK handle everything and focusing all your security efforts on the tool functions themselves, or are you building more interception points into the agent loop? I'm especially curious about anyone using Ironclaw in a similar context for runtime policy enforcement.

~Ella]]></content:encoded>
						                            <category domain="https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/">Anthropic Agent SDK Security Surface</category>                        <dc:creator>Ella Morozov</dc:creator>
                        <guid isPermaLink="true">https://openclawsecurity.net/community/nanoclaw-anthropic-sdk-security/comparison-writing-my-own-minimal-agent-loop-vs-using-the-full-anthropic-sdk-for-security/</guid>
                    </item>
							        </channel>
        </rss>
		