OpenAI's framework sets tracks and deadlines for disclosing model misalignment, launching with 6 reports from RL training ...
Nunchux AI's training-free VC-Attention fixes value outliers and FP32 softmax bottleneck, speeding video DiT attention up to ...
Google Research's Retrieve-for-Train uses RL once to train a 53.9M-parameter diffusion retriever, delivering 12× to 20× ...
Stanford's Paper2Agent turns research papers into tested MCP servers, letting AI agents run their methods through natural ...
Prior Labs' TabPFN-3.5 beats the 2015 Otto winner with default settings and ranks first on 7 tabular benchmarks.
Google launches Gemini 3.8 Live and Extended Thinking, native speech to speech models with background tool calling, 97 ...
Agent-net open sources Webagent, a Go harness that builds guarded, multi-channel AI agents from one declarative JSON spec.
I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is ...
Reward AI's OM-1 is a general-purpose robot policy trained only on human demonstrations, running across arms and humanoids ...
claude plugin eval scores realistic prompts with 6 grader types; 4 are free, llm and baseline bill a judge model. Every case ...
Sakana AI's PC-ALM adds per-layer Lagrange multipliers to predictive coding, matching backprop and training 1000-layer networks locally.
Meta FAIR's AI Research Preference Models rank unexecuted ML candidates, raising AIRS-Bench from 0.684 to 0.729 without ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results