Co-founder and CEO of Rhesis AI. We build the layer that gets expert knowledge into AI agents.
Rhesis (MIT) is a shared workspace where domain experts annotate agent behavior and engineers get feedback they can actually act on.
Most AI quality tools assume your team already knows what to test. That assumption puts one engineer in the middle, collecting requirements from product, domain experts, and support, then translating all of it into evals. Rhesis starts a step earlier. The people who know the domain define what good looks like, and that definition becomes test cases the whole team reuses.
What that means in practice:
- One expert annotation expands into a set of test cases, and the original judgment stays attached
- Every review points at the case and the agent version it refers to, so the reasoning survives the release
- Domain experts work in the UI, engineers work through the SDK and MCP, both against the same record
- Runs alongside what you already use: DeepEval, RAGAS, Garak, PyRIT, and 60+ metrics
- Open source core, Enterprise Edition when you need SSO and RBAC
- Featured in the Thoughtworks Technology Radar (2026)
- Shipped v0.10: RBAC and SSO, redesigned Insights, Microsoft Agent Framework integration
- Running EvalOps Unfiltered, a Berlin meetup series on LLM evaluation in production
- PhD in digital entrepreneurship, research at the Hasso Plattner Institute
- 10+ years building and scaling digital products, apps and SaaS
- Former Managing Director at Navigating Art, an art information SaaS
- Former Managing Director at HPI Ventures, a pre-seed VC
- Shipping Rhesis from Potsdam, Germany
- Talking to teams about how they decide what good means for their agents, and writing down what I learn
- Growing the team and the open source community
Our mascot is a platypus. It hunts by picking up electrical signals from things buried in the mud, which is roughly what we do for agent behavior. Unicode has no platypus, so the duck is standing in. 🦆



