Agent and RAG evaluation

Build fast. Know what works.

I help agent and RAG teams evaluate, understand, and optimize the systems they are building—especially tool calling, skills, retrieval, routing, and model choices. I bring machine learning discipline to agent development so you can find the biggest lift, protect the weakest link, and scale on measured results.

Agent development is moving at software speed inside a non-deterministic, statistical world. That creates enormous opportunity—and dangerous blind spots.

The answer is not to slow down. It is to turn speed into productivity by making the system observable, testable, interpretable, and improvable.
What I do

Bring machine learning tools to your agent flow.

I work across diagnosis, experimental design, evaluation infrastructure, optimization, implementation, and technical education. The goal is practical: understand what is happening, identify leverage, and make the system materially better.

01 · Agent evaluation

Define success from evidence—not instinct.

Create modular, testable evaluation criteria grounded in how users and domain experts actually judge outcomes.

  • Golden datasets and contrastive cases
  • Modular LLM-judge rubrics
  • Human-to-LLM judgment alignment
  • Node-level and end-to-end scoring
02 · Tool and skill optimization

Give the agent a better surface to reason over.

Reduce ambiguity in tools, MCPs, and skills. Test naming, descriptions, grouping, masking, and exposure strategies.

  • Tool-choice and order analysis
  • Skill comparison and rewriting
  • Hierarchical tool exposure
  • Gateway and dependency mapping
03 · Harness optimization

Find the configuration that actually performs.

Match jobs to models, prompts, tools, skills, and retrieval strategies using structured experiments instead of noisy averages.

  • Model / prompt / tool factor testing
  • Best-worst and pairwise evaluation
  • Interaction-effect identification
  • Routing and deployment recommendations
Read my harness-optimization method →
04 · RAG and context systems

Make retrieval structured, interpretable, and useful.

Diagnose retrieval failures and shape growing knowledge into navigable, task-specific context rather than a flat pile of chunks.

  • Retrieval and answer diagnostics
  • Hierarchical and clustered search
  • Context routing and reduction
  • Second-brain and knowledge architecture
05 · Trace and drift analysis

See the branches your agents actually create.

Analyze request distributions, tool calls, transitions, outputs, failures, and changing usage patterns over time.

  • First / next / before / after analysis
  • Failure-mode and path clustering
  • Question and answer distribution drift
  • Operational dashboards and alerts
06 · ML for agent engineers

Build statistical intuition into the team.

Hands-on workshops and working sessions that make experimental design, uncertainty, classification, ranking, and optimization directly useful to agent developers.

  • Internal engineering tutorials
  • Evaluation-system design reviews
  • Statistical debugging practices
  • Reusable LangSmith evaluation patterns
The approach

Cultivate the wild growth.

AI is growing like a forest: fast, fertile, tangled, and full of opportunity. I help turn that growth into an orchard—structured enough to inspect and harvest, dynamic enough to keep adapting.

NURTURE

Strengthen what is already working.

Identify promising patterns, successful skills, reliable routes, and high-value use cases. Give them the evidence and structure needed to scale.

PRUNE

Remove ambiguity and wasted growth.

Retire weak configurations, redundant tools, misleading descriptions, noisy context, and branches that create cost without creating value.

HARVEST

Convert experimentation into results.

Deploy winners, improve routing, codify learning, and bring tools, skills, evals, and decisions back into a dependable operating system.

Engagement model

From pain point to measurable lift.

I can enter at the point where a system is unclear, unreliable, expensive, difficult to scale, or simply not living up to its promise.

Diagnose

Map the architecture, traces, decisions, metrics, and user expectations. Find the places where uncertainty is hiding.

Design

Create the smallest credible evaluation or experiment that can separate signal from noise and reveal leverage.

Implement

Build the evaluation pipeline, analysis, routing logic, data structure, dashboard, or workflow needed to make the insight operational.

Scale

Integrate the result into LangSmith or your existing stack, train the team, and establish a loop that keeps learning from production.

I learned this in the most volatile market in the world.

Knowledge Algorithms grew out of Energy Algorithms, and from the work of helping build one of the Midwest's most successful virtual electricity trading companies. The operating belief was simple: wild and chaotic data should be cultivated into enough structure to be accessible at any moment, while remaining dynamic enough to adapt to a changing market.

Electricity markets can move in extreme and unexpected directions. That volatility is not only risk—it is the source of opportunity. The work demanded quantitative discipline, fast feedback, strong judgment, and systems that could evolve without becoming opaque.

AI has the same shape of opportunity.

Agents and RAG systems are proliferating quickly. Tools, skills, prompts, retrievers, models, and orchestration layers are growing in every direction. The creativity is real. So is the chaos.

I do not want to reproduce the same old software patterns with an LLM attached. I want to work where the problems are difficult and the upside is massive: evaluation, context, tool use, routing, optimization, knowledge structure, and continual learning.

A problem solver who can move from idea to system.

My working rhythm is direct: assess risks, grab opportunities, iterate. I analyze pain points, verify operation, research alternatives, design experiments, implement solutions, and communicate clearly with technical and non-technical teams.

The infrastructure must be ready to pivot, identify the new thing, and remain robust under deadline. Data must be labeled, modeled, and structured correctly. Teams need visual observability—enough to know where they are without distraction, right now. A solid core is what makes creative, dynamic work possible.

I am especially interested in working with LangChain and LangSmith teams and customers, where small evaluation and orchestration ideas can become scalable infrastructure. The outcome I aim for is confidence that insight can be reached from every part of the data stack, whatever the system encounters.

Systems I am developing

Practical structures for growing agent systems.

These are not rigid product packages. They are recurring solution patterns that can be adapted to a company's data, architecture, and operating needs.

SKILL GARDEN

Cultivate reusable intelligence.

Gather skill versions, group and contrast them, evaluate them against real work, and rewrite them for the use case. Preserve the creativity and analytical power of the company in an evolving, testable space.

TOOL LATTICE

Structure the agent's action space.

Create a dynamic hierarchy over tools, MCPs, and workflows. Reduce context burden, expose clearer contrasts, control gateways, and track the paths agents follow.

VISUAL SEARCH

Make knowledge navigable.

Turn documents, tools, skills, and internal knowledge into hierarchical, visual structures that both people and agents can traverse without flattening everything into a ranked list.

Selected work

Ideas backed by working systems.

My projects explore the same underlying question from different angles: how can data and decisions be structured so that people and agents can act with less friction and more confidence?

Stop Ranking Agent Configs by Average Score

My practical harness-optimization method for separating model, prompt, and tool effects using best-worst evaluation instead of noisy averages.

Read on TDS →
Pairwise Cross-Variance Classification

Zero-shot embedding classification using cross-modal pairwise structure; improved text/image agreement from 61% to 89% in the published experiment.

Read on TDS →
Divariance

A research metric connecting Fréchet- and KL-like distributional differences for feature selection and interpretable embedding analysis.

Read paper →
VaultBubble

A browser research vault combining capture, semantic structure, topic maps, and retrieval for a more useful personal knowledge system.

Open project →
TreeSite & JoyRide

Hierarchical visual navigation and alternative controls for content-heavy and accessibility-sensitive experiences.

Ask for a demo →
Orthogonal Wonder

Essays and working ideas on machine learning, decision systems, agent evaluation, and useful structures for complex problems.

Read the Substack →

Build big on a real foundation.

You do not need less ambition. You need a clearer view of what is working, what is not, and where the next unit of effort creates the greatest lift.

Start a conversation