Coding benchmarks show what AI can produce. Firmulate’s live company test asks whether agents can manage pressure, finish crucial work and stay honest.
IBM unveils Granite 4.2, a family of dense, reasoning-focused language models in 3B, 8B, and 30B sizes, supporting tool calls and reinforcement learning.
Claude Will Now Watermark All Content Generated Using Its Tools – New Atlas
Anthropic announces that all AI-generated content from Claude will now include watermarks to improve detectability amid regulatory and industry pressure.
FDA’s Landmark Approval And What It Means For Consumer Health Safety
The FDA has approved a new targeted therapy for metastatic pancreatic cancer, marking a significant milestone in cancer treatment and impacting consumer health safety.
How Budget AI Is Driving Open-Weight Industry Disruption
Alibaba’s release of the open-weight Qwen3.8-Flash-Next model is reshaping AI distribution, emphasizing efficiency over raw performance amid rising Chinese model adoption.
Anthropic’s Claude AI experienced a widespread outage across web, mobile, and API platforms, but service has now been restored, impacting users and developers.
SenseTime Open-sources 8B Multimodal Model With Native 4K Image Output – TechNode
SenseTime has open-sourced an 8-billion-parameter multimodal AI model claiming native 4K image generation, with details on licensing and performance still pending.