AI News
Latest AI engineering news, updated daily.
Ai Engineering
NVIDIA Transfers 32K KV Caches 25x Faster
NVIDIA researchers mapped KV caches between matched language models in 278ms, cutting Qwen3 14B to 32B handoffs by 25.1x.
Nvidia · Kv Cache · Language Models
Ai Engineering
AVO Reaches 100% on ARC-AGI-3 Public Set
NVIDIA’s AVO harness solved all 183 ARC-AGI-3 public levels with Claude Opus 5, lifting performance from 30.16% to 100%.
Arc Agi · Nvidia Avo · Benchmark Results
Ai Engineering
GPT-5.6 Inference Spans 25 AWS Regions on Bedrock
AWS and OpenAI added cross-Region inference for GPT-5.6 Sol, Terra, and Luna across more than 25 Amazon Bedrock Regions.
Amazon Bedrock · Openai · Cross Region Inference
Ai Agents
Mistral Agentic Search lifts FinanceBench accuracy to 86%
Mistral AI released Agentic Search, an iterative retrieval layer that raises FinanceBench accuracy to 86% while reducing latency and token use.
Agentic Search · Retrieval Augmented Generation · Mistral Ai
Ai Agents
Binance Agent OS Gives AI Agents Access to Crypto Trading
Binance launched Agent OS, letting third-party AI agents analyze markets, access account data, and execute crypto trades within user-defined limits.
Ai Trading · Binance · Crypto Ai
Prompt Engineering
Encrypted prompts exposed Grok chat data in 40% of tests
Adversa AI disclosed an encrypted prompt injection that made Grok exfiltrate user metadata and chat history through its browsing tool.
Prompt Injection · Ai Security · Data Privacy
Ai Engineering
LFM2.5-DSpark Delivers 3.18x Faster Decoding on H100
Liquid AI released LFM2.5-DSpark draft models, reaching 3.18x faster decoding on H100 GPUs and reducing function-calling latency by 57%.
Inference Optimization · Speculative Decoding · H100 Gpus
Ai Engineering
CISA Adds MLflow CVE-2026-64849 to KEV Catalog
CISA added a critical MLflow SSRF vulnerability to its KEV catalog after attackers used it to target cloud metadata services and credentials.
Cybersecurity · Vulnerability Management · Mlflow
Ai Engineering
Liquid AI LFM2.5 Q4_0 Recovers 97% Accuracy via Distillation
Liquid AI has released new Q4_0 GGUF checkpoints for its LFM2.5 models using Quantization-Aware Distillation to retain up to 97 percent of BF16 accuracy.
Quantization · Distillation · Model Optimization
Ai Coding
GitHub Auth Retry Storm Triggers 7-Hour Global Outage
A failed component and subsequent authentication retry storm knocked GitHub core services and Copilot offline for over seven hours on August 17.
Github Outage · Github Copilot · Cloud Infrastructure
Ai Engineering
Firefox 154.0 Ships Zero-Retention AI Search via Exa
Mozilla released Firefox 154.0 with an optional Smart Window mode, integrating Exa-powered web retrieval and local visual history with zero data retention.
Browser Integration · Privacy Focused Ai · Web Retrieval
Ai Engineering
410 Tokens/Sec Nemotron 3.5 Lightning Hits SageMaker JumpStart
AWS has added NVIDIA's 30B Nemotron 3.5 Lightning model to SageMaker JumpStart, offering high-throughput speculative decoding for agent execution layers.
Nvidia Nemotron · Amazon Sagemaker · High Throughput Inference
Ai Engineering
First Jane Street Sohu Deployment Drives $21B Etched Valuation
Hardware startup Etched reached a $21 billion valuation following a $700 million Series D and the successful deployment of its first Sohu AI inference rack.
Ai Agents
IBM ALTK-Evolve Framework Drops Agent Memory Token Costs by 85%
IBM Research's new ALTK-Evolve framework optimizes AI agent memory injection by model tier, reducing token overhead by up to 85% while boosting task completion.
Memory Optimization · Token Efficiency · Ibm Research
Ai Coding
Cursor Origin launches bidirectional GitHub sync for AI agents
Cursor has launched Origin, a code-hosting platform designed for AI agents featuring bidirectional GitHub synchronization and native Vercel integrations.
Cursor Editor · Github Integration · Ai Software Development
Ai Engineering
Google PhotoScan Maps Insulin Resistance via Mobile Cameras
A deep learning model trained on 35,000 biobank scans allows standard smartphone cameras to estimate granular body composition and cardiometabolic risk.
Deep Learning · Healthcare Ai · Computer Vision
Ai Engineering
CoSnitch Exploit Leaks Copilot Data via Hidden URL Parameter
Varonis researchers disclosed CoSnitch, a vulnerability in Microsoft Copilot Personal that allowed one-click data exfiltration via a hidden autorun parameter.
Security Vulnerability · Microsoft Copilot · Data Exfiltration
Ai Agents
Warp's 6-Stage Agent Orchestration Layer Automates 30% of PRs
Warp Factories introduces a model-agnostic control plane that orchestrates fleets of AI agents across a six-stage software development pipeline.
Autonomous Agents · Software Engineering Automation · Agent Orchestration
Ai Agents
1,200 Custom Claude Agents Drop ABC Legal Deployment to 2 Hours
Anthropic's latest deployment reveals how ABC Legal transitioned its entire workforce to a decentralized builder model using Claude Managed Agents.
Anthropic Claude · Managed Agents · Workflow Automation
Ai Engineering
$350M Series A Completes Groq's Pivot to Nvidia Neocloud
Groq closed a $350 million Series A round at a $3.5 billion valuation, finalizing its transition from an AI chipmaker to an Nvidia-powered inference cloud.
Venture Capital · Cloud Infrastructure · Ai Hardware