Prompt & Optimization

Treating prompts and agent programs as something to be measured and improved automatically, rather than hand-tuned indefinitely.

5 projects

FaissVector & Storage

It is built around an index that holds a set of vectors and searches them, and the dozens of available index structures exist because the trade-offs are real: search time against result quality against memory per vector against how long the index takes to build, and whether it needs training data at all. Some methods keep only a compressed representation and never the original vectors, which is how a single server reaches billions of them. There is no server, no filtering and no access control here — those are what the databases built on top of it add.

OfficialMIT40.8k+26
DSPyOptimization

You declare what each step should do and supply a metric; the optimiser searches for the prompts and few-shot examples that maximise it. The payoff is that improving a pipeline becomes measurable rather than superstitious. It requires an evaluation set — without one there is nothing to optimise against.

OfficialMIT37.6k+156
OpikObservability

A decorator on a function or a wrapper around a client is enough to start recording traces, and from there the same data feeds experiments that score responses automatically instead of by hand — hallucination, answer relevance, context recall and the rest. More than thirty integrations cover the usual frameworks and providers, and the whole platform can run locally or on Kubernetes rather than only as a hosted service. Its difficulty is not capability but overlap: it competes directly with two platforms already in this catalogue, so the choice turns on integrations and hosting rather than on what is missing.

OfficialApache-2.021.7k+135
Agent LightningOptimization

The usual objection to reinforcement learning for agents is that it means rebuilding the agent as a training environment, at which point you are no longer improving the thing you run. Agent Lightning avoids that by putting a proxy where the model endpoint used to be: the agent keeps its own tools, prompts, control flow and environment, and every interaction that passes through the proxy becomes training data. A trainer updates the policy, a controller runs the agent locally or as Kubernetes jobs. Microsoft reports taking a 9B model from 41.8% to 56.4% on SWE-bench Verified using 6,000 samples, and publishes that pipeline. This is GPU work with a training stack behind it, not a library you add to an application.

OfficialMIT17.9k
GEPAOptimization

Reinforcement learning compresses an entire execution into one number, which throws away the part that says what went wrong. GEPA keeps it: error messages, profiling output and reasoning logs go back to a model that diagnoses the failure and proposes a specific edit, and it maintains a frontier of candidate prompts each strong on different examples rather than collapsing to one winner. The project reports reaching results in 100 to 500 evaluations where reinforcement learning needs upwards of 10,000. Both the running and the reflecting are model calls, so the efficiency is relative to RL, not to zero.

OfficialMIT6.3k+96