Laya runs offline on M4, ChatGPT tracks ads
New offline inference benchmarks, cross‑site data collection, and rescue services raise fresh questions for AI developers.
New offline inference benchmarks, cross‑site data collection, and rescue services raise fresh questions for AI developers.
NVIDIA launches four specialized inference platforms featuring the L4 and H100 NVL GPUs, aiming to speed generative AI workloads across cloud and enterprise.
OpenAI cuts GPT‑5.6 Sol fees by 50% while Moonshot AI launches Kimi K2.5 and K3, models that beat GPT‑5 on reasoning and bring trillion‑parameter open‑source to the market.
Kog challenges the notion that GPUs are ill-suited for agentic workflows, aiming to squeeze more inference out of them. The French startup's approach may change how we utilize GPUs.
Manifest shuts down its router while open‑source alternatives like RouteLLM, Any‑LLM, and ClawRouter vie for the cost‑saving niche.
Groq shifts focus to AI inference, raising $650M in funding. What's next for the chipmaker?