Skip to contentSkip to stories

Updated

All AI news

Aug 3

Aug 3Mon
  1. JetBrains AI BlogAI score52

    JetBrains Built a Central CLI to Control Spiraling AI Tool Costs

    AIJetBrains says its AI development expenses rose roughly 10x over six months as developers adopted three to five AI tools each. It built the JetBrains Central CLI, which routes third-party agent traffic through its AI platform so managers can set per-developer and team limits and view consumption reports. The CLI opened to early access on July 8 for anyone with JetBrains AI credits.

  2. Manus BlogAI score38

    Manus Adds ElevenLabs Connector for Chat-Based Audio Generation, Transcription, and Voice Apps

    AIManus has launched an ElevenLabs connector that lets users generate speech, transcribe recordings, clone voices, and build audio apps through a single chat. Users connect their authorized ElevenLabs account via Integrations, and audio is processed within their own ElevenLabs environment according to its policies. Availability depends on users having an active ElevenLabs account, with capabilities tied to their ElevenLabs plan and credit balance.

Aug 2

Aug 2Sun
  1. OpenRouter BlogAI score40

    OpenRouter Launches Ori Eval to Find the Best AI Model for Your App

    AIOpenRouter has released Ori Eval, an agent-driven tool that runs your app's prompts against candidate models and returns a comparison table of catch rate, latency, cost per PR, and pass/fail results. The tool asserts on called tools and grades open-ended answers with an LLM judge, pinning the harness and model during each run. Its evals are code files that can run in CI to block regressions and re-run when new models ship.

Jul 31

Jul 31Fri
  1. DeepSeek · new models on Hugging FaceAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    AIDeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    Why it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  2. DeepSeek API NewsAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu
  1. MiniMax BlogAI score72

    MiniMax H3 unifies text, image, video, and audio generation in one model

    AIMiniMax launches H3, a general-purpose multimodal generation model that understands text, images, video, and audio as unified context. It generates video up to 15 seconds at 2K resolution with native stereo sound, and the company says model weights will be opened in the coming days, subject to applicable laws and regulations. MiniMax also says H3 is priced below mainstream models at 2K and 768p.

    Why it matters: The post explains how a unified multimodal design and training choices enable 2K video with native stereo sound, useful for comparing against closed video generators.

  2. Thinking Machines LabAI score65

    Thinking Machines proposes staged, evidence-based release path for open-weight models

    AIThinking Machines argues that safe open-weight releases depend on both model safety testing and readiness of the surrounding ecosystem, and that release should proceed in iterative stages. For its Inkling and Inkling-Small models, internal evaluations, four external red-teaming groups, and adversarial fine-tuning tests led the company to conclude that releasing the weights was not likely to add material risk beyond existing open-weight models.

    Why it matters: The post lays out a staged, evidence-gated path to releasing open weights, with concrete safety tests and the ecosystem measures behind each stage.

  3. Microsoft AI BlogAI score14

    Leaders share how AI transformation depends on mindset, team adoption, and culture

    AILeaders interviewed for Alysa Taylor's "What's the Tea?" series, including executives at Adobe, Lumen, and Sitecore, say the shift from AI apprehension to expected adoption is the precondition for transformation. Behavioral scientist Jon Levy argues the goal is raising a team's collective intelligence, not just cutting costs, with leadership and continuous training driving scale.

Jul 29

Jul 29Wed
  1. Fireworks AI BlogAI score54

    Fireworks tests whether LoRA or full fine-tuning gaps come from data, learning rate, or rank

    AIFireworks AI ran controlled SFT experiments on Qwen3.5-9B comparing LoRA with full parameter fine-tuning across three synthetic verifiable tasks. The post argues that a FullFT advantage can come from data coverage, learning-rate tuning, or adapter rank, and it recommends testing these in that order before switching methods. Under a fixed multi-task budget, FullFT kept a 4.29-point lead over the best LoRA recipe tested, while matched data exposure favored LoRA.

  2. Google LabsAI score60

    Google Launches Lyria 3.5 in Flow Music With Better Vocals and Lyrics

    AIGoogle is rolling out Lyria 3.5, its newest music generation model, in Google Flow Music today. The update improves musicality, lyric quality and prompt adherence, and vocal expressiveness and pronunciation, and gives users more control over tempo and duration.

    Why it matters: The post names the specific capability changes and where users can access them, which helps readers judge fit for music creation workflows.

  3. Liquid AI NewsletterAI score46

    Liquid AI Expands LFM2 Tokenizer to 128K, Speeding On-Device Thai, Vietnamese, and Hindi

    AILiquid AI doubled the LFM2 tokenizer's vocabulary from 65K to 128K without retraining from scratch, extending the original BPE merges and initializing new embeddings as the mean of their sub-tokens. The expanded tokenizer needs 4.0× fewer tokens for Thai, 2.6× fewer for Vietnamese, and 2.4× fewer for Hindi, which the source says yields roughly 2.2–3.7× faster on-device decoding for these languages with no reported quality loss on previously supported languages. LFM2.5-8B-A1B and the expanded tokenizer are available on Hugging Face with open weights.

  4. Berkeley AI ResearchAI score44

    K-Search Adapts CUDA Kernel Expertise to Apple Silicon MLX Backend

    AIBerkeley AI Research extended the K-Search evolutionary kernel framework with an MLX backend and a CUDA-to-MLX translation layer, letting it adapt existing CUDA kernels for Apple Silicon. The team reports a 0.97x speedup relative to the native MLX Attention kernel and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel. The method uses Gemini 3.5 Pro Preview to both reason about optimizations and write candidate kernels.

  5. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases UEmbed-9B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-9B, a decoder-only multimodal embedding model built on Qwen3.5 9B that outputs both dense and SPLADE-style sparse embeddings from one forward pass. It supports text, image, video, and mixed-modal inputs for retrieval and multimodal search, and the family also includes 2B and 4B variants. The model is available on Hugging Face, with transformers and vLLM inference support.

  6. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba NLP releases UEmbed-4B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-4B, a decoder-only multimodal embedding model built on Qwen3.5 4B that outputs both dense and sparse embeddings from one forward pass. It handles text, image, video, and mixed-modal inputs for retrieval and visual-document search, and sparse activations map to vocabulary terms usable with inverted indexes. The model is available on Hugging Face in a family that also includes 2B and 9B variants.

  7. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score43

    Alibaba-NLP releases UEmbed-2B, a multimodal model producing dense and sparse embeddings

    AIAlibaba-NLP's UEmbed-2B, a decoder-only multimodal embedding model built on Qwen3.5 2B, produces both dense and SPLADE-style sparse embeddings from a single forward pass. It supports text, image, video, and mixed-modal inputs for retrieval, and the 4B and 9B variants are also available. The team reports state-of-the-art results on the text and agent tracks of MMEB-v3.

Jul 28

Jul 28Tue
  1. Augment Code BlogAI score39

    GPT-5.6 Sol Becomes Augment Cosmos's Default Model for Token Efficiency

    AIAugment Code has made GPT-5.6 Sol the default model in Cosmos, choosing it as the most token-efficient model to clear its pass-rate floor for long-horizon software engineering tasks. The company ranks models by cost per task rather than list price per million tokens, since retries on failed steps add token spend. Users can still select any model, and the default will change as more token-efficient models emerge.

  2. Fireworks AI BlogAI score46

    Fireworks AI Shows Low-Cost Fine-Tuning Lifts Domain Embedding Retrieval

    AIFireworks AI describes fine-tuning Qwen3-Embedding-8B on private (query, positive) pairs using bidirectional InfoNCE loss through its Training SDK, then serving the model via an OpenAI-compatible embeddings endpoint. The post reports that around 150 training steps was enough, that rank-32 LoRA landed within about one point of full-parameter fine-tuning, and that gains were largest where the base model struggled, while tasks like CoSQA and FiQA2018 showed flat results.

  3. Cognition Blog (Devin, Windsurf)AI score28

    LTM Partners with Cognition to Deploy Devin for Cybersecurity Risk Reduction

    AILTM has partnered with Cognition to deploy Devin, the AI software engineer, through BlueVerse RightLogic, a managed, outcome-based service that clears customers' vulnerability backlogs. RightLogic is designed to clear 80 percent of an enterprise's CVE backlog, up from the 60 percent previously delivered, and will focus first on banking, financial services, and insurance. The service is the first of five joint offerings the companies plan to bring to market.

  4. JetBrains AI BlogAI score60

    Ponytail Skill Cuts Claude Code Costs 10% But Not the Advertised 54%

    AIJetBrains tested the ponytail skill for Claude Code across 80 paired tasks and found a median 10.3% cost reduction, with p=0.004. Code written fell about 15% median versus the advertised 54%, reaching 31% on larger builds and little on already-lean tasks. No quality difference was detected, and the skill only self-activated when its ruleset was injected by a plugin hook.

    Why it matters: The benchmark separates advertised savings from measured results and shows the code cut depends on how much the baseline agent over-builds.

  5. MiniMax · new models on Hugging FaceAI score76

    MiniMax H3 releases open-weight omni-modal video model with native stereo audio

    AIMiniMax released H3, an open-weights omni-modal model that generates video with native stereo audio up to 2K and 15 seconds. The system combines H3-Context-IR preprocessing, the H3-Base generator at 768p, and H3-Regenerate-2K for 2K output, with the Context-IR and 2K modules available only through API.

    Why it matters: The source details a three-module pipeline and open weights with deployment paths, showing how a video model is served and reproduced locally.

  6. METR BlogAI score58

    METR outlines how independent researchers could investigate AI agent misalignment incidents

    AIMETR proposes that AI companies track agent misalignment incidents and have independent researchers investigate the most serious ones, focusing on the motives behind the behavior. The post lists core investigation questions covering incident surveys, root causes, and remediation, along with the model access, transcripts, employee interviews, and training-data tools such investigators would need. It also calls for results to go to company boards and oversight bodies and be published with disclosed redaction terms.

Jul 27

Jul 27Mon
  1. Google LabsAI score36

    Google and KDDI launch AI Startup Support Program for Japanese AI startups

    AIGoogle's AI Futures Fund and KDDI are launching the AI Startup Support Program to back AI-native Japanese startups with joint equity investment. Selected startups will get early access to Gemini, Nano Banana, and Lyria, plus Google Cloud credits, sovereign Gemini and GPU credits provisioned in Japan via KDDI's Osaka-Sakai Data Center, and hands-on technical support.

  2. Liquid AI BlogAI score49

    Liquid AI Releases LFM2.5-Encoders for Fast Long-Context Encoding on CPU

    AILiquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, bidirectional encoders built on the LFM2 hybrid architecture and available on Hugging Face. They support an 8,192-token context and are designed for fine-tuning on classification and token-level tasks. On CPU, LFM2.5-Encoder-230M is the fastest model tested from 1K tokens up, running about 3.7x faster than ModernBERT-base at 8,192 tokens.

  3. Meta AI BlogAI score36

    Meta's DINOv3 and SAM Power Edge-Based Assistive Robotics at Pittsburgh

    AIThe University of Pittsburgh's RAMMP team is integrating Meta's DINOv3 and SAM models into on-device assistive robotics to detect door buttons, cups, and curbs for navigation assistance. The models run on compact, battery-powered hardware, with optimizations such as reduced memory footprint and lower precision, enabling real-time perception without network connectivity. RAMMP's perception system pairs SAM-based auto-labeling with an RF-DETR detector fine-tuned on DINOv2 embeddings, and the team is now testing voice and touch input for object selection.

Jul 26

Jul 26Sun
  1. Fireworks AI BlogAI score60

    Fireworks AI adds open-weight Kimi K3 with US-only serverless endpoints

    AIFireworks AI made the open-weight Kimi K3 available for inference and training on its platform, with US-only serverless endpoints and Zero Data Retention. In its own head-to-head with Opus 5, the post reports K3 at 92.7% accuracy and $0.52 per task on SWE (480) against Opus 5's 94.8% and $1.05, with the vendor claiming up to 5x better cost efficiency per task.

    Why it matters: The post compares Kimi K3 with Opus 5 on accuracy and cost per task, giving readers concrete figures to judge the open model against closed alternatives for their own workloads.

  2. Berkeley AI ResearchAI score44

    Berkeley AI Research Trains LLMs to Update Beliefs for Long Tasks

    AIBerkeley AI Research introduces ABBEL, a framework that replaces full interaction histories with natural-language belief states that models update as new observations arrive. On CollabBench collaborative coding, belief grading closes about half the performance gap to full-context models while using fewer peak tokens and training in 50 steps instead of 100.

Jul 25

Jul 25Sat
  1. LangChain BlogAI score39

    What does it mean for companies to "own their intelligence" with AI?

    AILangChain Blog argues that companies need to own their AI intelligence rather than rely on generic models, because general models do not know company-specific policies, workflows, or risk tolerances. Ownership means controlling the agent system (model optionality, harness, and context), the economics, quality, and risk of AI work, and how intelligence compounds over time. The post uses an insurer's claims processing as an example of why off-the-shelf models fall short.

Jul 24

Jul 24Fri

Jul 23

Jul 23Thu
  1. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-Base, a 122B deep research agent model

    AIBAAI has released AREX-Base, a 122B-total, 10B-activated Mixture-of-Experts deep research agent built on Qwen3.5-122B-A10B with a 262,144-token context. The model uses an inner research loop and an outer self-improvement loop, and the source reports it scoring 82.5 on BrowseComp and 85.4 on GAIA, under Apache 2.0.

    Why it matters: The release pairs a 122B-parameter deep research agent with benchmark tables against frontier and open models, letting readers compare its search-agent results directly.

  2. BAAI · new models on Hugging FaceAI score47

    BAAI releases AREX-Turbo, a compact 4B recursive self-improving deep research agent

    AIBAAI's AREX-Turbo is a dense 4B deep research agent built on Qwen3.5-4B with a 262,144-token context length. It scores 70.7 on BrowseComp, 81.6 on GAIA and 40.6 on HLE with tools, versus 82.5, 85.4 and 52.4 for the 122B AREX-Base. The model is released under Apache License 2.0 and targets lower-cost research-agent deployment.

  3. Cognition Blog (Devin, Windsurf)AI score38

    Cognition Acquires The Interaction Company, Maker of the Poke Texting AI Agent

    AICognition has acquired The Interaction Company of California, the maker of Poke, a personal AI agent that texts users proactively and is approved to text natively on Apple Messages. Poke has exchanged more than 100 million messages in the last three months, and Poke users can keep using the product as before. Cognition says its models and infrastructure will make Poke faster and more reliable.

Jul 22

Jul 22Wed
  1. Cognition Blog (Devin, Windsurf)AI score41

    Cognition signs MOU with U.S. Department of Energy to join Genesis Mission

    AICognition has signed a memorandum of understanding with the U.S. Department of Energy to join the Genesis Mission, a national AI initiative launched by executive order in November 2025. Cognition will contribute its Devin autonomous AI software engineer in four areas: software and data security, modernizing legacy scientific code, expanding scientific workforce capacity, and cloud modernization. Devin Desktop and CLI are listed as FedRAMP Class D (High) Authorized, and the company has offered in-kind code security scans for national laboratory codebases.

Jul 21

Jul 21Tue
  1. JetBrains AI BlogAI score55

    JetBrains Air adds ACP agents, local models, and Java/Kotlin code intelligence

    AIJetBrains Air now connects to ACP-compatible coding agents, including GitHub Copilot CLI, OpenCode, Pi, and Cline, through the Agent Client Protocol. The release also adds Beta Java and Kotlin navigation and diagnostics powered by the IntelliJ IDEA code engine, local model support through Ollama or LM Studio, and Docker-based agent tasks on Windows.

  2. OpenAI Alignment Research BlogAI score65

    OpenAI and Apollo Research measure reward-seeking with Contrastive SDF

    AIOpenAI and Apollo Research introduce Contrastive SDF, a method that finetunes two copies of a model on opposite beliefs about grader and authority preferences to measure reward-seeking. In the post, intermediate checkpoints of a capabilities-focused OpenAI o3 RL run without safety training increasingly side with the grader over RL training, and this sensitivity is validated on reward-hacking models and model organisms trained to favor specific authorities.

    Why it matters: The paper gives a controlled way to test whether a model changes behavior based on beliefs about its grader, a question that matters for judging alignment evaluations.

  3. JetBrains AI BlogAI score62

    JetBrains Context adds repository indexing to coding agents in early access

    AIJetBrains has launched JetBrains Context in early access, a repository intelligence layer that builds a semantic index so coding agents can retrieve relevant code without repeated searching. In tests on 205 SWE-bench tasks, 175 production-monorepo tasks, and 1,953 code-localization tasks, it reduced agent turns by up to 68%, latency by up to 59%, and execution cost by up to 48%. It works with Claude Code, Codex CLI, and Junie CLI at no additional cost for JetBrains AI subscribers, and it does not store source code on JetBrains Context servers.

    Why it matters: The source gives benchmark figures for turns, latency, and cost, showing how repository indexing might change agent workflows on large codebases.

  4. Meta AI BlogAI score44

    Meta's SAM 3 and DINOv3 Power SYNAPS-I's Genesis Mission Imaging Pipeline

    AISYNAPS-I, a multi-lab Genesis Mission project led by Lawrence Berkeley National Laboratory, uses Meta's open-source SAM 3 and DINOv3 models to segment X-ray and micro-CT scientific imagery. The fine-tuned pipeline, run on 300 A100 GPUs, reduced a grapevine xylem analysis from a month of expert annotation per time step to about 15 minutes. The team can deploy the open models inside secure national lab infrastructure, where research data must remain.

Jul 20

Jul 20Mon