# Loamist Evals

> Agents that hold 98%+, run after run. Reliability for enterprise agents.

Canonical: https://www.loamist.com/evals

Loamist builds evals and nested evals that make agents work inside your enterprise. Our techniques beat variance, reach your reliability threshold, and hold 98%+ accuracy consistently.

- **98%+**: Accuracy, run after run
- **Nested**: Multi-step evals
- **Threshold**: Set before we build
- **In prod**: Enterprise scale

## Fig. 01 · Same task, run after run: Variance is the problem. Nested evals remove it.

An illustrative dot-column chart of repeated runs of the same task against a 98% threshold. Each eval depth narrows the run-to-run spread.

| Eval depth | Mean accuracy | Spread, run to run |
|---|---|---|
| Single agent (off-the-shelf model) | 81.2% | ±11 pts |
| + Evals (checks on every output) | 90.4% | ±6 pts |
| + Nested evals (evals of the evals) | 96.1% | ±2.6 pts |
| + Customer-tuned preferences (Loamist knowledge base graph) | 99.0% | ±0.5 pts |

## § 01 · What makes it work: Reliability is measured, not hoped for.

1. **We know how to build a rigorous eval system.** We define the threshold a decision has to reach, then build the evals that prove it, before the agent ships.
2. **Multi-step evals are required; we build them.** Evals, and evals of those evals, that find and remove variance, not just errors.
3. **In prod, at enterprise scale.** 98%+ held across repeated runs on live work, not one good result on a good day.

## Start

Have a problem that needs reliability? Trust us with it.

- [Talk to the evals team](https://www.loamist.com/contact.html)
- [About Loamist](https://www.loamist.com/)

Loamist products: [Loamist Trade Finance](https://www.loamist.com/trade-finance) · [Loamist Geospatial](https://www.loamist.com/geospatial) · [Loamist Evals](https://www.loamist.com/evals) · [Loamist Services](https://www.loamist.com/services)
