Skip to content

#Agent

Oct 8

TodayOct 8Thu16 items
  1. Xiaomi MiMoAI score63

    Xiaomi releases MiMo-V2.5-TTS series of speech synthesis models

    Xiaomi released the MiMo-V2.5-TTS Series, three speech synthesis models for stock voices, voice design, and voice cloning. The models accept natural-language style instructions and inline audio tags, and the source says the three models are free of charge for a limited time on the Xiaomi MiMo API platform. Xiaomi also open-sourced integration Skills for agent applications on GitHub.

    AIWhy it matters: The release shows how a TTS family adds style instructions, inline audio tags, and voice design or cloning to speech synthesis, which matters for agent and creative workflows.

  2. ClaudeDevsAI score42

    ICYMI we've cut the price of Sonnet 5.5 cache reads in half. In the Claude Platform, it's now $0.10 per million tokens (input is $2, output is $10). This means Sonnet 5.5 now runs ~20% cheaper on most agentic work. (API only, no change to Claude Code usage limits.)

    ICYMI we've cut the price of Sonnet 5.5 cache reads in half. In the Claude Platform, it's now $0.10 per million tokens (input is $2, output is $10). This means Sonnet 5.5 now runs ~20% cheaper on most agentic work. (API only, no change to Claude Code usage limits.)

  3. Artificial AnalysisAI score62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    Artificial Analysis added trusted-access models to its Cyber Index, and GPT-6 Sol (Daybreak Blue, max) now ranks first. The model is available only through OpenAI's Daybreak program and records no safety blocks across the Index. Its overall score is 32 points higher than the publicly available GPT-6 Sol (max), at a cost of $1.77 per task versus $11.67 for Grok 4.7 (xhigh).

  4. OdysseyAI score38

    Odyssey-3 is a foundation world model, enabling many applications in physical AI, human experiences, and even how we train intelligences. We're particularly excited by agents learning from experience inside Odyssey-3, working to accomplish objectives.

    Odyssey-3 is a foundation world model, enabling many applications in physical AI, human experiences, and even how we train intelligences. We're particularly excited by agents learning from experience inside Odyssey-3, working to accomplish objectives.

  5. MarkTechPostAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    JetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  6. Leiphone (雷峰网)AI score62

    Claude Haiku 5.5 gains on computer use but still trails Sonnet 5.5 in terminal coding

    Anthropic released Claude Haiku 5.5, raising its OSWorld 2.1 score from 15.7% to 72.4% and supporting a 1 million token context window. The article notes Haiku 5.5 still scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, and that prompts above 100,000 tokens are priced higher, so migration costs need to be measured on real workloads.

  7. OpenRouterAI score44

    Step 5 Preview from @StepFun_ai is live on OpenRouter. Their new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it: https://openrouter.ai/stepfun/step-5-preview

    Step 5 Preview from @StepFun_ai is live on OpenRouter. Their new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it: https://openrouter.ai/stepfun/step-5-preview

  8. JetBrains AI BlogAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    JetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    AIWhy it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

  9. The DecoderAI score72

    Claude Haiku 5.5 cuts prices but uses more tokens than GPT-6 Luna

    Anthropic released Claude Haiku 5.5, its fastest and most affordable small model, at prices up to 90 percent lower for most prompts under 100,000 tokens. Artificial Analysis ranks it first among small-class models on its Intelligence Index with a score of 43, but it consumes about three times the output tokens per task that GPT-6 Luna needs at maximum effort.

Oct 7

Oct 7Wed
  1. IThome · AI (IT之家)AI score72

    Anthropic releases Claude Haiku 5.5, cutting run costs about 75% from Haiku 4.5

    Anthropic released Claude Haiku 5.5, which it calls the fastest, cheapest, and most capable Haiku model so far. On average it costs about 75% less to run than Haiku 4.5, with input at $0.10 and output at $0.50 per million tokens for requests up to 100,000 tokens. Anthropic also cut Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says lowers run costs by about 20% on many agent tasks.

  2. MarkTechPostAI score67

    Anthropic releases Claude Haiku 5.5, a small model with 1M context

    Anthropic has released Claude Haiku 5.5, its cheapest and fastest small model, priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. It keeps a 1M token context window, up to 128K output tokens, and is generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic reports 72.4% on OSWorld 2.1 (offline subset) versus 15.7% for Haiku 4.5, and the article notes that non-default temperature, top_p or top_k values return a 400 error.

  3. AWS Machine Learning BlogAI score56

    Claude Haiku 5.5 becomes available on Amazon Bedrock and Claude Platform on AWS

    Anthropic's Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic, it is the fastest and most efficient model in the Claude 5.5 family and costs around 75 percent less than Claude Haiku 4.5 for most tasks. The post also covers pairing it with Claude Opus 5.5 as a subagent layer and provides Boto3, Converse, and Anthropic SDK examples for calling the model.

  4. ThariqAI score67

    Claude Haiku 5.5 returns as a cheaper, faster small model

    Anthropic has released Claude Haiku 5.5, which it describes as the cheapest, fastest, and most capable small model it has released. On average it costs around 75% less to run than Claude Haiku 4.5, and the author says it is 10x cheaper than Haiku 4.5 under 100k tokens. It can be tried with computer use, workflows, and the API.

    This story has a top pick“Anthropic releases Claude Haiku 5.5, scoring 43 on the Intelligence Index”

  5. Testing CatalogAI score62

    Anthropic releases Claude Haiku 5.5, its fastest and cheapest model

    Anthropic has released Claude Haiku 5.5, which the author describes as its fastest and cheapest model to date. The source says it costs about 75% less to run than Claude Haiku 4.5 and is the first Haiku model with an adjustable effort setting. The attached benchmark table reports Haiku 5.5 scores on tasks including computer use (OSWorld 2.1 offline subset, 72.4%) and Terminal-Bench 4.0 (39.2%), compared with Haiku 4.5 and other models.

  6. ClaudeAI score38

    Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles repetitive work like summaries and classification, and pairs well with Claude Opus 5.5 and Sonnet 5.5 as a subagent on coding work. It’s also fast enough for live customer support and browser use.

    Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles repetitive work like summaries and classification, and pairs well with Claude Opus 5.5 and Sonnet 5.5 as a subagent on coding work. It’s also fast enough for live customer support and browser use.

  7. Liquid AIAI score23

    d1-3B ranks first among models under 10B on the Decision Index v0.2.1, a benchmark for structured decision-making. Built from LFM2.5-VL-3B, it makes decisions from text and images in one pass. Use it for reranking, agent guardrails, and visual inspection. 2/

    d1-3B ranks first among models under 10B on the Decision Index v0.2.1, a benchmark for structured decision-making. Built from LFM2.5-VL-3B, it makes decisions from text and images in one pass. Use it for reranking, agent guardrails, and visual inspection. 2/

Oct 6

Oct 6Tue
  1. Guillaume LampleAI score42

    On important coding and agentic benchmarks such as DeepSWE, AutomationBench, and AA-Briefcase, ML4 matches the performance of the best open-weight models. It is SOTA on finance and legal workflows, as well as on complex multimodal grounding benchmarks. It can navigate complex terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks. 4/n

    On important coding and agentic benchmarks such as DeepSWE, AutomationBench, and AA-Briefcase, ML4 matches the performance of the best open-weight models. It is SOTA on finance and legal workflows, as well as on complex multimodal grounding benchmarks. It can navigate complex terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks. 4/n

  2. Latent SpaceAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    Reflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.

Oct 5

Oct 5Mon
  1. IThome · AI (IT之家)AI score49

    Reflection AI releases open-weight Beam model to rival DeepSeek and Kimi

    Reflection AI, an Nvidia-backed startup, released Beam, its first open-weight large model, aimed at coding and agent tasks. The company says Beam is comparable to Z.ai's GLM-5.2 and is approaching Qwen3.8-Max on coding and agent work. Beam has 501 billion total parameters, with 23 billion activated per task in a sparse architecture.

  2. Sophia YangAI score62

    Reflection AI's Beam open model has 501B total parameters and 23B active

    Sophia Yang congratulated Reflection AI on Beam, a 501B-parameter open model with 23B active per token. She attributes its efficiency to an RL length penalty that discourages unnecessary tokens and a sparse MoE architecture. Reflection says full weights will be released this month, and the quoted post reports training over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks.

  3. Clément DelangueAI score72

    Reflection AI announces Beam, a 501B-parameter agentic open model

    Reflection AI introduced Beam, an agentic open model with 501B total parameters and 23B active parameters, trained end-to-end from scratch. The quoted announcement says it targets frontier reasoning efficiency and coding and agentic tasks, with full weights due this month. Clément Delangue, Hugging Face's CEO, reposted it with a welcome to the Reflection organization on Hugging Face.

    AIWhy it matters: The quoted announcement names Beam's parameter scale, active-parameter count, and coding and agentic focus, which helps readers gauge where it fits among open models.

Oct 2

Oct 2Fri
  1. Hugging Face BlogAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    Ai2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    AIWhy it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

Oct 1

Oct 1Thu

Sep 30

Sep 30Wed
  1. indigoAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    Google has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    AIWhy it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.