Prompt Regression CheckGuide

Moving OpenAI Evals graders to promptfoo: the mapping, grader by grader

Updated 2026-10-11 · by Hieu Tran, written with AI agents and checked against the sources below

OpenAI Evals become read-only on 31 October 2026 and the API shuts down on 30 November 2026. OpenAI points to promptfoo as a destination, but its guide has you recreate tests and assertions by hand. Most graders have a direct promptfoo equivalent; a few need care because the scores aren't on the same scale.

Free tool: Export all your evals in your browser: bundle JSON, dataset JSONL and a promptfoo config

Paste an API key; the page calls api.openai.com directly from your browser and gives you the files. Nothing goes through our servers.

The mapping

Templates

OpenAI templates use {{item.field}} for dataset columns and {{sample.output_text}} for the model's answer. In promptfoo, columns are plain {{field}} under each test's vars, and the answer is output (in a rubric, {{output}}).

Keep the old results

A promptfoo run is a new run: scores from a different judge prompt or library won't match OpenAI's to the decimal. Export your past runs and output items as well as the test cases, so you can compare the first promptfoo run with the last OpenAI one and spot graders that changed meaning in the move.

Free tool: Export all your evals in your browser: bundle JSON, dataset JSONL and a promptfoo config

Paste an API key; the page calls api.openai.com directly from your browser and gives you the files. Nothing goes through our servers.

Which of your graders was hardest to move, and what did you do with your past run results?

We're researching this problem and read every answer. Tell us what happened (4 short questions, no sign-up; AI tools help us read the answers).

More guides