Researchers from Fudan, Tencent Hunyuan, SJTU and Shanghai AI Lab publish WorldReward, a chunk-level VLM reward that jointly scores ...
The NLRB counted ballots Sept 3 in the Wiki Workers United U.S.-CWA recognition election, with staff voting 158-14 to unionize after the ...
Gimlet Labs, which builds software that splits AI inference workloads across different chip architectures (GPUs and SRAM-centric silicon), raised $300M ...
OpenEvidence launched three production medical AI models for verified clinicians: Osler (~5s responses), Sackett (~30s), and Snow (~5min for ...
A new arXiv preprint by Zixuan Fu, Bingxiang He and eleven co-authors argues that on-policy distillation of large language models can be pushed to most ...
Puffin-World, a new unified multimodal model from Nanyang Technological University's S-Lab and collaborators including the University of Michigan ...
Follow-up study finds one-shot on-policy distillation — training on a single query — recovers ~71.5% of full-dataset state coverage within 100 steps ...
OpenAI has pledged $1 billion in subsidized access to its Daybreak cyber models, training and technical support for organizations defending water, ...
A paper posted to arxiv this week proposes training terminal agents against environments that grow harder "generation by generation" instead of ...
"Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe ...
Researchers introduce a training method that incrementally evolves terminal-agent environments off-policy generation by generation, validated ...
A new paper argues KV-cache eviction doesn't need sophisticated token-importance scoring: preserving the prompt and evicting uniformly at random ...