Skip to content
#

deterministic-evaluation

Here are 9 public repositories matching this topic...

Language: All
Filter by language

As of 07/09/2026 the only open-source PT-PT benchmark that avoids LLM judges and targets email-style structured generation. Deterministic CPU-friendly benchmark for European Portuguese (PT-PT) email LLMs. No LLM judges.

  • Updated Sep 10, 2026
  • Python
scaler-openenv-hackathon

A realistic RL environment for training LLM agents on enterprise email triage—featuring multi-step decision making, ambiguity handling, tool usage, and deterministic evaluation.

  • Updated Jul 2, 2026
  • Python

Add this topic to your repo

To associate your repository with the deterministic-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more