Forum

Notifications
Clear all

TIL: How to use eBPF to trace self-hosted agent calls

5 Posts
5 Users
0 Reactions
28 Views
(@agent_trace_runner)
Eminent Member
Joined: 3 months ago
Posts: 18
Topic starter   [#1522]

I've been auditing a self-hosted agent runtime built on LangChain, and the central question was always data flow: where do prompts, retrieved context, and final outputs actually travel? The vendor-hosted dashboards show you the sanitized, high-level ops, but for a real security assessment, you need kernel-level visibility into the execution path, especially when those agents are making external API calls.

This is where eBPF becomes indispensable. By instrumenting the syscall layer, you can trace the entire lifecycle of an agent's execution without modifying the application code. The goal was to capture a concrete trace of a self-hosted agent performing a web search and processing the result. Here's a simplified version of the eBPF program (using BCC for Python bindings) that hooks `execve` and `connect` to map the process lineage and network calls:

```python
from bcc import BPF
import ctypes as ct

bpf_text = """
#include
#include

struct data_t {
u32 pid;
u32 ppid;
char comm[TASK_COMM_LEN];
char argv[256];
};
BPF_PERF_OUTPUT(events);

int trace_execve(struct pt_regs *ctx) {
struct data_t data = {};
data.pid = bpf_get_current_pid_tgid() >> 32;
data.ppid = bpf_get_current_ppid();
bpf_get_current_comm(&data.comm, sizeof(data.comm));
bpf_probe_read_user_str(&data.argv, sizeof(data.argv), (void *)PT_REGS_PARM2(ctx));
events.perf_submit(ctx, &data, sizeof(data));
return 0;
}
"""

# ... (Additional probes for connect syscall to track outbound HTTP requests)
```

Running this while the agent operates reveals the chain: the main Python interpreter spawning subprocesses, and crucially, the outbound TCP connections to external APIs (like Serper or a vector database). The trace output shows the exact command lines and destination IP:port pairs. This is the kind of granularity you lose with vendor-hosted solutions—their observability stack will tell you an "LLM call" happened, but not the specific syscall sequence or the raw socket-level data before encryption.

The tradeoff is clear. Self-hosting shifts the operational burden and security responsibility onto your team, but in return, you gain the ability to deploy this level of deep inspection. You can answer critical questions: Was the retrieved context from the web search exfiltrated to an unexpected endpoint? Did the agent process spawn any unexpected child processes? With vendor-hosted runtimes, you're relying on their logging pipeline, which is often abstracted and filtered for "privacy" or "simplicity." For high-stakes deployments involving sensitive data, that black-box nature is a non-starter. The ability to trace with eBPF is a compelling argument for accepting the self-hosting burden, provided you have the expertise to manage and interpret the output.



   
Quote
(@quinn_mod2)
Eminent Member
Joined: 3 months ago
Posts: 16
 

That's a solid approach for mapping process execution. I'd be careful about assuming the connect syscall gives you the full picture on API calls though. A lot of these frameworks use higher-level HTTP libraries that might not trigger a straightforward connect for every request, especially with connection pooling or if they're going through a local proxy.

Have you considered also hooking into SSL/TLS write events? Seeing the encrypted stream is one thing, but if you can correlate it with the process lineage from execve, you get a much clearer idea of what data is actually leaving the box, even if it's wrapped.


/q


   
ReplyQuote
(@vendor_skeptic)
Eminent Member
Joined: 3 months ago
Posts: 22
 

You're right, connect alone is a noisy mess for this. SSL/TLS writes get you closer, but now you're in the arms race of BPF program complexity vs. library implementations.

For a security audit, I'd skip trying to reconstruct the plaintext in-kernel and just capture the socket descriptor and target IP/port from connect, then grab the process memory containing the plaintext before the SSL_write. Attach a uprobe to the library's actual send function. Less elegant, but you get the actual data payload, not just metadata.

Otherwise you're just tracing that something happened, not what.


show me the proof, not the whitepaper


   
ReplyQuote
(@red_team_learner_ivy)
Eminent Member
Joined: 3 months ago
Posts: 22
 

This is exactly the kind of low-level visibility I'm trying to get. How reliable is this for tracing the agent's own internal steps, like when it calls a tool? Does hooking `execve` catch subprocesses spawned by something like LangChain's `ShellTool`, or does that get lost?


Breaking things to learn.


   
ReplyQuote
(@eve_redteam)
Eminent Member
Joined: 3 months ago
Posts: 24
 

Tracing execve is a decent start, but you're already trusting the runtime too much. That BPF snippet will show you the subprocess spawned by a ShellTool, sure. But what about when the agent's Python interpreter itself calls `requests.post` or uses `ctypes` to load a shared library that makes the call? You'll see the parent python process, not the actual code path.

The real blind spot is that these frameworks are moving towards in-process tool calling, specifically to avoid the overhead of fork/exec. If your audit hinges on seeing a new process appear, you'll miss the entire class of vulnerabilities where the agent manipulates its own runtime memory to, say, redirect an API call.

You need to pair this with uprobes on the framework's own dispatch functions, or you're just watching the front door while the data slips out a window.


reality has a bias against your threat model


   
ReplyQuote