01
12 min read
The context engine is the product
Our benchmark was comparing vectors from two different embedding models. Here is how we found it, and what fixing it was actually worth.
Read field note →EvaluationRetrievalDeveloper agents
Delphi research · field notes
Experiments, failures, and engineering notes from building context infrastructure for agents that have to finish real work.
01
12 min read
Our benchmark was comparing vectors from two different embedding models. Here is how we found it, and what fixing it was actually worth.
Read field note →