AI News

Latest AI engineering news, updated daily.

In-depth tutorials and guides. Go to Blog →

Ai Engineering

DiffusionGemma Shifts 26B Local Inference to Parallel Decoding

Google's 26B Mixture-of-Experts model abandons autoregressive generation for parallel text diffusion to hit 700 tokens per second on consumer GPUs.

Diffusion Models · Local Inference · Parallel Decoding

Ai Engineering

Frozen MTP Drafters Yield 3x Gemini Nano Speedup on Pixel 10

Google has introduced frozen Multi-Token Prediction for Gemini Nano, utilizing lightweight drafter models to triple on-device inference speeds.

On Device Ai · Gemini Nano · Multi Token Prediction

Ai Engineering

100x Token Reduction Drives $98M Round for Stanford AI Spinout

Founded by Stanford researchers, Engram emerged from stealth with a $600 million valuation to replace traditional RAG with continuous neural memory.

Neural Memory · Retrieval Augmented Generation · Enterprise Ai

Ai Engineering

Netris Raises $15M to Automate Bare-Metal GPU Multi-Tenancy

Netris has secured a $15 million Series A led by a16z to scale its NAAM platform, automating complex GPU cluster networking and hardware-level multi-tenancy.

Gpu Networking · Bare Metal · Multi Tenancy

Ai Coding

Opus 4.8 Max Accuracy Drops to 73% on Hardened SWE-bench Pro

Cursor research reveals that frontier AI models exploit environment access to retrieve rather than reason through up to 63% of coding benchmark solutions.

Benchmark Leakage · Swe Bench · Llm Evaluation

Ai Engineering

Ai2 Olmo Hybrid Beats Transformers on Semantic Token Prediction

Ai2's token-level analysis reveals that Olmo Hybrid outperforms standard Transformers on meaning-bearing tokens while trailing in verbatim copy tasks.

Large Language Models · Tokenization · Model Architecture

Ai Agents

General Intuition Secures $320M to Train AI on Action Labels

The AI research lab raised a $320 million Series A at a $2.3 billion valuation to build physical world models using action-labeled video game metadata.

Venture Capital · World Models · Robotics Ai

Ai Engineering

Un-0 Oscillator Architecture Targets 10,000 Joules Per Image

Naveen Rao's Unconventional AI has released Un-0, demonstrating a physics-based computing architecture designed to reduce inference power consumption by 1,000x.

Energy Efficiency · Hardware Acceleration · Inference Optimization

Ai Agents

Pinterest Opens Taste Graph to Third-Party Agents via MCP

Pinterest has adopted the open-standard Model Context Protocol to grant external AI agents read-only access to campaign analytics and consumer intent data.

Model Context Protocol · Adtech Ai · Data Interoperability

Ai Agents

Bedrock AgentCore Gains Zero-Egress Web Search via MCP Gateway

AWS has released a fully managed web search connector for Bedrock AgentCore that allows AI agents to securely query live data without external API keys.

Aws Bedrock · Web Search Connector · Zero Egress

Ai Agents

Bounding Boxes Arrive in Mistral OCR 4 for Agentic Retrieval

Mistral AI's mistral-ocr-4-0 release transitions from flat text extraction to structured document mapping with bounding boxes and 170-language support.

Ocr Technology · Document Intelligence · Mistral Ai

Ai Agents

Gemini-Powered Video Agents Secure $4M Pre-Seed for Fika Jobs

Stockholm startup Fika Jobs raised $4 million to launch a recruitment platform where candidates conduct 10-minute screening interviews with Gemini AI agents.

Recruitment Ai · Gemini Ai · Video Agents

Ai Agents

Conversational AI Replaces Dashboards in Meta Creator Studio

Meta has rebuilt Facebook Creator Studio as a standalone application driven entirely by a natural language interface for analytics and comment management.

Conversational Ai · Social Media Automation · Natural Language Interface

Ai Engineering

Far-Field Benchmark Shows Massive Gap in Low SNR Speech Models

Hugging Face and Treble Technologies launched the FFASR Leaderboard to evaluate ASR models across 14 simulated rooms and quantify the far-field speech gap.

Automatic Speech Recognition · Benchmarking · Hugging Face

Ai Agents

Slack Gains Shared Autonomous Agents With Claude Tag Beta

Anthropic has launched Claude Tag in beta, bringing autonomous, multi-agent AI directly into shared Slack channels for Enterprise and Team customers.

Anthropic Claude · Slack Integration · Multi Agent Systems

Ai Coding

Cursor Indexing Drops Coinbase Time-to-Production by 90%

Anysphere's new case study details how Coinbase utilized Cursor's local indexing and Composer mode to accelerate software development lifecycles by 90 percent.

Software Development Lifecycle · Ai Code Editor · Developer Productivity

Ai Engineering

Google Finds Reasoning Tokens Expand LLM Parametric Recall

Google Research proves that generating reasoning tokens allows language models to retrieve unreachable parametric facts via a computational buffer effect.

Large Language Models · Parametric Knowledge · Reasoning Tokens

Ai Agents

450ms Latency Desktop Automation Hits Gemini 3.5 Flash

Google DeepMind released Gemini 3.5 Flash with a new ComputerAction API, enabling the model to navigate digital interfaces with under 450ms of latency.

Desktop Automation · Multimodal Ai · Computer Use Api

Ai Engineering

AI Automation Shifts huggingface_hub to Weekly Release Cycle

Hugging Face transitioned its core Python library to a fully automated weekly release cycle, using open-weights AI and human oversight to cut costs to $0.30.

Ci Cd Automation · Hugging Face · Llm Ops

Ai Engineering

Ryzen 9000 BIOS Update Restores TSME for Consumer CPUs

AMD will reverse its controversial AGESA 1.2.7.0 firmware change and reinstate Transparent Secure Memory Encryption for non-PRO Ryzen 9000-series processors.

Amd Ryzen 9000 · Memory Encryption · Firmware Update