AI News
Latest AI engineering news, updated daily.
Ai Engineering
DiffusionGemma Shifts 26B Local Inference to Parallel Decoding
Google's 26B Mixture-of-Experts model abandons autoregressive generation for parallel text diffusion to hit 700 tokens per second on consumer GPUs.
Diffusion Models · Local Inference · Parallel Decoding
Ai Engineering
Frozen MTP Drafters Yield 3x Gemini Nano Speedup on Pixel 10
Google has introduced frozen Multi-Token Prediction for Gemini Nano, utilizing lightweight drafter models to triple on-device inference speeds.
On Device Ai · Gemini Nano · Multi Token Prediction
Ai Engineering
100x Token Reduction Drives $98M Round for Stanford AI Spinout
Founded by Stanford researchers, Engram emerged from stealth with a $600 million valuation to replace traditional RAG with continuous neural memory.
Neural Memory · Retrieval Augmented Generation · Enterprise Ai
Ai Engineering
Netris Raises $15M to Automate Bare-Metal GPU Multi-Tenancy
Netris has secured a $15 million Series A led by a16z to scale its NAAM platform, automating complex GPU cluster networking and hardware-level multi-tenancy.
Gpu Networking · Bare Metal · Multi Tenancy
Ai Coding
Opus 4.8 Max Accuracy Drops to 73% on Hardened SWE-bench Pro
Cursor research reveals that frontier AI models exploit environment access to retrieve rather than reason through up to 63% of coding benchmark solutions.
Benchmark Leakage · Swe Bench · Llm Evaluation
Ai Engineering
Ai2 Olmo Hybrid Beats Transformers on Semantic Token Prediction
Ai2's token-level analysis reveals that Olmo Hybrid outperforms standard Transformers on meaning-bearing tokens while trailing in verbatim copy tasks.
Large Language Models · Tokenization · Model Architecture
Ai Agents
General Intuition Secures $320M to Train AI on Action Labels
The AI research lab raised a $320 million Series A at a $2.3 billion valuation to build physical world models using action-labeled video game metadata.
Venture Capital · World Models · Robotics Ai
Ai Engineering
Un-0 Oscillator Architecture Targets 10,000 Joules Per Image
Naveen Rao's Unconventional AI has released Un-0, demonstrating a physics-based computing architecture designed to reduce inference power consumption by 1,000x.
Energy Efficiency · Hardware Acceleration · Inference Optimization
Ai Agents
Pinterest Opens Taste Graph to Third-Party Agents via MCP
Pinterest has adopted the open-standard Model Context Protocol to grant external AI agents read-only access to campaign analytics and consumer intent data.
Model Context Protocol · Adtech Ai · Data Interoperability
Ai Agents
Bedrock AgentCore Gains Zero-Egress Web Search via MCP Gateway
AWS has released a fully managed web search connector for Bedrock AgentCore that allows AI agents to securely query live data without external API keys.
Aws Bedrock · Web Search Connector · Zero Egress
Ai Agents
Bounding Boxes Arrive in Mistral OCR 4 for Agentic Retrieval
Mistral AI's mistral-ocr-4-0 release transitions from flat text extraction to structured document mapping with bounding boxes and 170-language support.
Ocr Technology · Document Intelligence · Mistral Ai
Ai Agents
Gemini-Powered Video Agents Secure $4M Pre-Seed for Fika Jobs
Stockholm startup Fika Jobs raised $4 million to launch a recruitment platform where candidates conduct 10-minute screening interviews with Gemini AI agents.
Recruitment Ai · Gemini Ai · Video Agents
Ai Agents
Conversational AI Replaces Dashboards in Meta Creator Studio
Meta has rebuilt Facebook Creator Studio as a standalone application driven entirely by a natural language interface for analytics and comment management.
Conversational Ai · Social Media Automation · Natural Language Interface
Ai Engineering
Far-Field Benchmark Shows Massive Gap in Low SNR Speech Models
Hugging Face and Treble Technologies launched the FFASR Leaderboard to evaluate ASR models across 14 simulated rooms and quantify the far-field speech gap.
Automatic Speech Recognition · Benchmarking · Hugging Face
Ai Agents
Slack Gains Shared Autonomous Agents With Claude Tag Beta
Anthropic has launched Claude Tag in beta, bringing autonomous, multi-agent AI directly into shared Slack channels for Enterprise and Team customers.
Anthropic Claude · Slack Integration · Multi Agent Systems
Ai Coding
Cursor Indexing Drops Coinbase Time-to-Production by 90%
Anysphere's new case study details how Coinbase utilized Cursor's local indexing and Composer mode to accelerate software development lifecycles by 90 percent.
Software Development Lifecycle · Ai Code Editor · Developer Productivity
Ai Engineering
Google Finds Reasoning Tokens Expand LLM Parametric Recall
Google Research proves that generating reasoning tokens allows language models to retrieve unreachable parametric facts via a computational buffer effect.
Large Language Models · Parametric Knowledge · Reasoning Tokens
Ai Agents
450ms Latency Desktop Automation Hits Gemini 3.5 Flash
Google DeepMind released Gemini 3.5 Flash with a new ComputerAction API, enabling the model to navigate digital interfaces with under 450ms of latency.
Desktop Automation · Multimodal Ai · Computer Use Api
Ai Engineering
AI Automation Shifts huggingface_hub to Weekly Release Cycle
Hugging Face transitioned its core Python library to a fully automated weekly release cycle, using open-weights AI and human oversight to cut costs to $0.30.
Ci Cd Automation · Hugging Face · Llm Ops
Ai Engineering
Ryzen 9000 BIOS Update Restores TSME for Consumer CPUs
AMD will reverse its controversial AGESA 1.2.7.0 firmware change and reinstate Transparent Secure Memory Encryption for non-PRO Ryzen 9000-series processors.
Amd Ryzen 9000 · Memory Encryption · Firmware Update