Vibrant Labs Research
Search
Search
Dark mode
Light mode
Reader mode
Explorer
Home
❯
sources
sources
68 items under this folder.
Sep 19, 2026
SoK: DARPA's AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned
Sep 19, 2026
Auditing Reward Hackability in Code RL Training Environments
Sep 19, 2026
EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
Sep 19, 2026
Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
Sep 19, 2026
HiL-Bench (Human-in-Loop Benchmark) Do Agents Know When to Ask for Help?
Sep 19, 2026
Rate or Fate? RLVεR: Reinforcement Learning with Verifiable Noisy Rewards
Sep 19, 2026
SWE-Mutation: Can LLMs Generate Reliable Test Suites in Software Engineering?
Sep 19, 2026
τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction
Sep 19, 2026
The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?
Sep 19, 2026
When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR
Sep 18, 2026
Agent² RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
Sep 18, 2026
Can Generalist Agents Automate Data Curation?
Sep 18, 2026
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
Sep 18, 2026
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
Sep 18, 2026
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Sep 18, 2026
τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Sep 10, 2026
Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP
Sep 10, 2026
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Sep 09, 2026
Automated Benchmark Auditing for AI Agents and Large Language Models
Sep 09, 2026
Agent Mentor: Framing Agent Knowledge through Semantic Trajectory Analysis
Sep 09, 2026
Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems
Sep 09, 2026
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
Sep 09, 2026
AgentSimulator: An Agent-based Approach for Data-driven Business Process Simulation
Sep 09, 2026
An Optimisation Framework for Unsupervised Environment Design
Sep 09, 2026
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
Sep 09, 2026
Autocurricula and the Emergence of Innovation from Social Interaction: A Manifesto for Multi-Agent Intelligence Research
Sep 09, 2026
BERTScore: Evaluating Text Generation with BERT
Sep 09, 2026
Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation
Sep 09, 2026
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
Sep 09, 2026
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
Sep 09, 2026
ClawEnvKit: Automatic Environment Generation for Claw-Like Agents
Sep 09, 2026
Constitutional AI: Harmlessness from AI Feedback
Sep 09, 2026
Designing Benchmarks for Knowledge Work
Sep 09, 2026
Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design
Sep 09, 2026
Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
Sep 09, 2026
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
Sep 09, 2026
Evolving Curricula with Regret-Based Environment Design
Sep 09, 2026
FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets
Sep 09, 2026
From Event Logs to Governed Action: A BlueSky Agenda for Agentic Process Mining
Sep 09, 2026
Holistic Evaluation of Language Models
Sep 09, 2026
ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
Sep 09, 2026
Training language models to follow instructions with human feedback
Sep 09, 2026
LOGIGEN: Logic-Driven Generation of Verifiable Agentic Tasks
Sep 09, 2026
Log analysis is necessary for credible evaluation of AI agents
Sep 09, 2026
MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents
Sep 09, 2026
More Than Reading Comprehension: A Survey on Datasets and Metrics of Textual Question Answering
Sep 09, 2026
Open-Endedness is Essential for Artificial Superhuman Intelligence
Sep 09, 2026
Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions
Sep 09, 2026
Prioritized Level Replay
Sep 09, 2026
QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization
Sep 09, 2026
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
Sep 09, 2026
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Sep 09, 2026
Replay-Guided Adversarial Environment Design
Sep 09, 2026
SABER: Small Actions, Big Errors — Safeguarding Mutating Steps in LLM Agents
Sep 09, 2026
SAGE: A Top-Down Bottom-Up Knowledge-Grounded User Simulator for Multi-turn AGent Evaluation
Sep 09, 2026
STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios
Sep 09, 2026
Scaling Laws for Neural Language Models
Sep 09, 2026
Towards a Science of AI Agent Reliability
Sep 09, 2026
Simulated Customers Never Walk Away: Decision Fidelity of LLM User Simulators Measured Against Real Purchase Outcomes
Sep 09, 2026
SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization
Sep 09, 2026
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
Sep 09, 2026
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
Sep 09, 2026
Training Compute-Optimal Large Language Models
Sep 09, 2026
VISTA: A Versatile Interactive User Simulation Toolkit for Agent Evaluation
Sep 09, 2026
WebArena: A Realistic Web Environment for Building Autonomous Agents
Sep 09, 2026
WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks
Sep 09, 2026
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Sep 09, 2026
World of Workflows: A Benchmark for Bringing World Models to Enterprise Systems