Amplifying.ai benchmarks 2,430 Claude Code runs, revealing strong preferences for custom DIY implementations and modern stack tools over traditional defaults.
Research WireIBM Research and UC Berkeley used the MAST taxonomy on ITBench to analyze 310 agent traces, finding GPT-OSS-120B failed 87.6% of tasks while frontier model Gemini-3-Flash reached 75.5% recall.
Research WireResearchers have unveiled *-PLUIE, a novel evaluation metric for AI text. By measuring an LLM's confidence on 'Yes/No' answers, it offers a faster, more computationally efficient alternative to traditional LLM-as-a-judge methods.
Research WireDoes AI fine-tuning impart new skills or just unlock existing knowledge? A new research paper proposes 'task complexity' as a formal metric to finally settle the debate.
Research WireResearchers present DMTS-NC, an AI method achieving up to 5.64x speedups in molecular dynamics simulations using non-conservative forces and neural potentials.
Research WireA new paper accepted at COLT 2026 establishes convergence bounds for discrete diffusion models, proving an iteration complexity order of O(d/ε).
Research WireA new research paper introduces Pep, a framework achieving 80.8% preference alignment while requiring 3-5x fewer interactions than standard RL models.
Research WireA study by Xiaosheng Zhao and team shows simple neural networks trained on legacy telescope data generalize effectively to modern, high-resolution surveys.
Research WireMacroGuide uses persistent homology to steer molecular diffusion models, raising macrocycle molecule generation rates from 1% to 99% for drug discovery.
Research WireResearchers fine-tuned the Poseidon PDE model using 13 GPU hours to emulate Martian weather, achieving a 34.4% performance boost on atmospheric forecasts.
Research WireA new paper expands Geometric Deep Learning to orbifolds—complex spaces with unique 'folded' symmetries. This breakthrough enables AI to analyze complex structured data.
Research WireMIT and partner researchers introduced canonical diffusion models that set state-of-the-art benchmarks on GEOM-DRUG without equivariant constraints.
Research Wire