Repositories list
64 repositories
RewardHarness
PublicSelf-evolving agentic reward framework for image-editing evaluation — 47.4% on EditReward-Bench from only 100 preference demos, no reward-model training. arXiv …VLM2Vec
PublicThis repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]ClawBench
PublicOpen-source benchmark for browser AI agents on daily tasks.Context-Forcing
PublicPixel-Reasoner
PublicPixel-Level Reasoning Model trained with RL [NeuIPS25]StructEval
PublicEvaluating LLMs' abilities to generate structural output [TMLR2025]EditReward
PublicEditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing [ICLR 2026]verl-tool
PublicFIM-Midtraining
PublicSWE-QA-Pro
PublicOpenResearcher
PublicOpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory SynthesisRationalRewards
PublicVideoEval-Pro
PublicVideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation [TMLR26]SWE-Next
PublicVisPhyWorld
PublicCritique-Coder
PublicHierarchical-Reasoner
PublicImagenWorld
PublicStress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks [ICLR 2026]MMLU-Pro
PublicThe code and data for "MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark" [NeurIPS 2024]EvolveCoder
PublicMMMU
PublicVisualWebInstruct
PublicVisCoder2
PublicThe official code of "VisCoder2: Building Multi-Language Visualization Coding Agents" [ICLR26]BrowserAgent
PublicMantis
PublicVideoScore2
PublicVideoScore
Publicofficial repo for "VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation" [EMNLP2024]ImagenHub
PublicA one-stop library to standardize the inference and evaluation of all the conditional image generation models. [ICLR 2024]General-Reasoner
PublicQuickCodec
Public
ProTip! When viewing an organization's repositories, you can use the
props. filter to filter by custom property.