USK

02Operating companyUSK Group LLC

Gerfach Labs

Security research and authorized assessment of AI agents and the tool surfaces they expose.

Founder-led, with the same people on every engagement.

644

Tests

24

Modules

6

Protocols

Hundreds

Endpoints evaluated

Dozens

MCP servers

01What Gerfach studies

The layer between an agent and the world.

Agents now call tools, tools call services, and protocols such as MCP have become the connective layer between them. A protocol contract says what a tool is allowed to do. Runtime behavior says what it actually does. Gerfach studies the gap.

The current body of work covers 644 tests across 24 modules and 6 protocols, evaluating hundreds of endpoints and dozens of MCP servers.

One mark per test.

  • 01

    Over-privileged tools

    Tools that can do far more than the task in front of them, and agents that will use every capability they are handed.

  • 02

    Tool injection

    Instructions that arrive through tool descriptions, results and retrieved content rather than from the user.

  • 03

    Credential exposure

    Keys, tokens and session material that leak through tool metadata, logs, error paths and agent memory.

  • 04

    Unsafe execution

    Tool surfaces that execute, fetch or write with no isolation between the model's intent and the host's authority.

  • 05

    Authorization boundaries

    Where a boundary is declared in a schema but does not hold at runtime, and who is accountable for the action that follows.

  • 06

    Reproducibility and auditability

    Whether a finding can be replayed by the owner, on their own infrastructure, with the same result.

02How an assessment runs

Declared behavior first. Runtime behavior second. Evidence always.

  1. 00

    Scope and ingest

    Targets, time windows, prohibited actions, tool descriptors, schemas and agent cards are normalized into one authorized surface map before any probe runs.

    Output

    Scope letter and surface map

  2. 01

    Analyze declared behavior

    Passive modules inspect names, annotations, auth boundaries, schema drift and capability claims without invoking side effects.

    Output

    Passive finding set

  3. 02

    Probe runtime behavior

    Active probes run only inside an isolated, network-sealed sandbox. Every outbound callback or file write is captured and tied back to its probe by a unique canary token.

    Output

    Canary trace and side-effect log

  4. 03

    Package the evidence

    Captured signals are linked from entry point through capability to impact, scored, and delivered as a replayable proof bundle for the technical owner.

    Output

    finding.yaml, repro.sh, scorecard

03What a finding contains

Findings without reproducible evidence are not findings.

finding.yaml

The finding itself: scope, preconditions, affected surface, severity rationale.

repro.sh

A script the owner runs on their own infrastructure to reproduce the result.

canary trace

Captured callbacks and side effects, each tied to the probe that caused it.

scorecard

Risk scored across fixed axes so two findings can be compared honestly.

Every finding is delivered as a replayable bundle. The owner runs it on their own infrastructure and gets the same result, or the finding is withdrawn.

04Research programs

Research that improves the assessment.

  • RA-01

    Browser and computer-use agent security

    Agents that operate real user interfaces. Visual prompt injection, DOM-based exfiltration, screen-grab side channels and intent hijacking.

  • RA-02

    Retrieval and vector store security

    Embedding inversion, retrieval poisoning and namespace bleed in vector databases. The new supply chain for context-augmented agents.

  • RA-03

    Agent-to-agent protocol security

    Trust transitivity in multi-agent systems. When agent A calls agent B, who is responsible for the action that follows?

  • RA-04

    Agent forensics and session replay

    When an autonomous agent acts incorrectly in production, how do you reconstruct what happened? Observability for non-deterministic systems.

Index of work

Published and forthcoming, as listed on gerfach.com. Coordinated disclosure with a 90-day vendor timeline.

  1. 01Across the Protocol Stack: An Empirical Security Analysis of AI Agent Tool Interfaces at ScaleForthcoming manuscript2026
  2. 02BrowserAgent Adversarial Benchmark (BAB-1)Forthcoming2026
  3. 03A2A Threat Model v1Forthcoming2026
  4. 04Coordinated disclosure policyPolicy2026
  5. 05On the difference between protocol contracts and runtime behaviorEssay, forthcoming2026
  6. 06Memory poisoning practitioner briefNote2026
  7. 07Capability attestation framework paperForthcoming2026
05Working with Gerfach

Authorized assessments begin with one assessable surface.

An MCP server, an agent, a tool surface. Scope is agreed in writing before any probe runs, and the result is a bundle the owner can replay. Founder-led, with the same people on every engagement.