Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
-
Updated
Aug 14, 2026 - Python
Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
[NeurIPS 2025] AGI-Elo: How Far Are We From Mastering A Task?
Benchmark methodology, task sets, and evaluation results for RA²R
A reproducible LiveCodeBench evaluation scaffold for controlled studies of code-reasoning SFT on Qwen2.5-1.5B.
Interactive benchmark for evaluating LLMs on clarifying ambiguous code requirements. Paper: arXiv:2607.00711
Self-Improving Agent for Code Generation
LoRA fine-tuning evaluation pipeline for code generation models on LiveCodeBench benchmark
“U.S. Provisional Patent Application No. 63/978,753, filed February 9, 2026” Aleph Generalizable Intelligence Convergence Engine (AGICE) - Evidence-Native, Policy-Governed Reasoning with Rollback and Geometrical Multi-Dimensional Transformation (GMDT) via the Complex Operator (−i)
Executable-code benchmark and verifier-guided post-training sandbox in the L20 project family.
NeoSmith Maestro — frontier-level coding accuracy from a bouquet of small models at ~1/30th the cost. Technical report + reproducible evidence (LiveCodeBench v6 92.2%, SWE-bench Pro 74.2/88.9).
To associate your repository with the livecodebench topic, visit your repo's landing page and select "manage topics."