
Observe
Instrument agent runs and MCP tool calls with OpenTelemetry, so every call has a name, a duration, an outcome and trace context, in the backend you already use.
I make AI agent and MCP pipelines observable and reproducible.
Describe your pipelineit got slower
it cost more
it behaved differently
This is for engineering teams with an agent or MCP-based system in production and a problem of exactly that shape: it failed, got slow, cost more or changed behaviour, and the logs do not say which step, which tool call or which prompt.

Instrument agent runs and MCP tool calls with OpenTelemetry, so every call has a name, a duration, an outcome and trace context, in the backend you already use.
Separate what the model decides from what code should decide. Fixed logic moves out of prompts into testable steps, so a run can be replayed, diffed and reviewed.

Design MCP servers and tool layers with explicit boundaries: read-only by construction, allowlisted calls, secrets kept out of the model's context.
POST /api/Orders/SendOrderrejected: not a GETGET /api/Orders/SendOrderrejected: path not in the setGET /api/Account/GetHoldingsSummaryrejected: the set holds the server's own spelling, SummeryGET /api/Account/GetHoldingsSummeryexact match: passesSession.request() raises unless the method is GET and the path is in one exact 15-entry allowlist. Paths shown are real; timing is illustrative.Five public, open-source repositories, all my own. Limits are listed where the READMEs list them.
Transparent Go proxy for MCP servers: one OpenTelemetry span per tools/call, no client or server changes. Speaks stdio, Streamable HTTP and HTTP+SSE.
Limit On HTTP+SSE, a server that advertises an absolute POST endpoint bypasses the proxy and you get zero tool-call spans. The README says so.
Keeps API keys out of a coding agent's context. Secrets live in the macOS Keychain; the agent only sees sealref:// references, and injected values are scrubbed from command output.
Limit macOS only: it is built on the Keychain.
Read-only MCP server over a brokerage API. The read-only guarantee is enforced at the session layer with an exact-GET allowlist, not by convention.
Single-binary agent harness in Rust: pluggable LLM providers, sub-agents, a skill library, and episodic memory in SQLite.
Curated list of tracing, evaluation, guardrail, gateway and MCP tooling for agent observability. Every entry checked to resolve and be maintained as of the last audit.
I do not publish a price list; scope decides the shape, and I quote after the first call.
Using the form below: what runs, what goes wrong, what you already use for monitoring.
Or an email thread, to confirm whether this is something I can help with. If it is not, I will say so and, where I can, point you elsewhere.
In writing: a time-boxed review with findings, a defined implementation task, or part-time contract work. You see the scope and rate before anything starts.
The more specific, the more useful my reply. Submitting sends these details straight to me; if that fails, a pre-filled email opens in your mail app instead.
I reply to every genuine inquiry by email.