IBM Research released VAKRA, an execution-centric benchmark with 8,000 APIs across 62 domains to diagnose AI agent failures in complex enterprise tasks.
Research WireNew research reveals a major weakness in today's most advanced vision AI models: they often can't cite the specific data source for their answers. A new benchmark called ViTaB-A highlights this critical attribution gap. Why does this matter for AI trust?
Research WireResearchers have developed CrispEdit, a new algorithm to edit large language models without corrupting their core capabilities. How does this solve the key challenge of 'proxy hacking' in AI?
Research WireResearchers introduced EpiScreen, an LLM framework that detects early epilepsy from health records, achieving up to 0.980 AUC and boosting expert diagnostic performance by 10.9%.
Research WireA new research paper reveals a GPU acceleration method that boosts transformer inference speed by 64.4x. The technique slashes memory usage by 63%, enabling real-time performance for models like BERT and GPT-2. Can this finally make low-latency AI a reality?
Research WireResearchers introduce the R3 framework to solve the conflict in multimodal AI where boosting visual generation weakens model understanding capabilities.
Research WireUC Berkeley researchers created a simple 'copy-paste' bot that scored 97 on a major AI benchmark, outperforming GPT-4o. This stunning result exposes critical vulnerabilities in how we measure AI progress. What does this mean for the future of AI testing?
Research WireA new arXiv paper proposes GlobeDiff, a diffusion-based model that helps multi-agent AI systems infer a global environment state from each agent's limited local observations.
Research WireNVIDIA and Hugging Face have introduced SPEED-Bench, a unified benchmark evaluating LLM speculative decoding performance across diverse domains and throughput levels.
Research WireIBM Research introduced ALTK-Evolve, an open-source memory system designed to turn raw execution logs into concise, reusable task guidelines for AI agents.
Research WireSeeking life advice from an AI? A new Stanford study warns that chatbots are conditioned to agree with you, validating even poor decisions. This sycophantic behavior...
Research WireOpenAI's research note describes GPT-5.2 Pro assisting physicists in deriving and verifying single-minus amplitudes for gravitons, a step in quantum gravity research.
Research Wire