OpenAI Evals become read-only on 31 October 2026 and the API shuts down on 30 November 2026. OpenAI points to promptfoo as a destination, but its guide has you recreate tests and assertions by hand. Most graders have a direct promptfoo equivalent; a few need care because the scores aren't on the same scale.
Free tool: Export all your evals in your browser: bundle JSON, dataset JSONL and a promptfoo config
Paste an API key; the page calls api.openai.com directly from your browser and gives you the files. Nothing goes through our servers.
string_check with eq → equals; ne → not-equals; like → contains; ilike → icontains. OpenAI's docs list the not-equal operation as both ne and neq, so check which your evals use.text_similarity depends on evaluation_metric: bleu → bleu, gleu → gleu, meteor → meteor, rouge_l → rouge-l, rouge_1…rouge_5 → rouge-n, cosine → similar (embeddings). Keep pass_threshold as threshold: both pass at or above it.fuzzy_match has no direct equivalent: promptfoo's levenshtein counts edits (pass at or below a number), not a 0–1 ratio. Port it as a small python or javascript assertion.score_model → llm-rubric with the grader's prompt as the rubric. promptfoo's judge scores 0–1, while OpenAI's range can be anything (say 1–10), so normalise: threshold = (pass_threshold − min) / (max − min). A pass at 7 on 1–10 becomes 0.667.label_model → llm-rubric, with the passing labels written into the rubric ("pass if the label is polite").python → promptfoo python assertion. OpenAI calls grade(sample, item) and expects a float; promptfoo calls get_assert(output, context), with your row in context['vars'], and accepts a bool, a float or a dict with pass, score and reason. Rename the arguments and read sample['output_text'] as output.multi graders are for reinforcement fine-tuning only, so they're not in Evals you need to move.OpenAI templates use {{item.field}} for dataset columns and {{sample.output_text}} for the model's answer. In promptfoo, columns are plain {{field}} under each test's vars, and the answer is output (in a rubric, {{output}}).
A promptfoo run is a new run: scores from a different judge prompt or library won't match OpenAI's to the decimal. Export your past runs and output items as well as the test cases, so you can compare the first promptfoo run with the last OpenAI one and spot graders that changed meaning in the move.
Free tool: Export all your evals in your browser: bundle JSON, dataset JSONL and a promptfoo config
Paste an API key; the page calls api.openai.com directly from your browser and gives you the files. Nothing goes through our servers.
Which of your graders was hardest to move, and what did you do with your past run results?
We're researching this problem and read every answer. Tell us what happened (4 short questions, no sign-up; AI tools help us read the answers).