Skip to contentSkip to stories

Updated

#Model release

Showing low-relevance items too. Hide low-relevance items

Sep 20

Sep 20Sun
  1. Qwen · new models on Hugging FaceOfficialAI score62

    Qwen releases open-source Qwen-Image-2.1 with a prompt rewriting model

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with a 7B-parameter visual generation component. The release also includes Qwen-Image-2.1-PE-T2I, a fine-tuned Qwen3.5-VL 9B model that rewrites brief image requests in any language into detailed English prompts with a recommended aspect ratio.

    Why it matters: The release pairs a 7B visual generation component with a separate prompt rewriting model, showing how a brief image request becomes a detailed English prompt before rendering.

  2. WanOfficialAI score10

    Wan posts motion control and animation demo clips

    AIWan (@Alibaba_Wan) promoted its motion control and animation capabilities with a short post, calling the results solid. The post includes no specifications, benchmarks, or availability details. A related Wan 3.0 experiment shared by another account shows a combat sport motion demonstration.

Sep 19

Sep 19Sat
  1. StepFunOfficialAI score29

    StepFun opens Step 5 Preview for trial on its platform

    AIStepFun invites developers to try the Step 5 Preview through its online platform. The post links to platform documentation for the model and a Discord community for discussion. No specific capabilities, benchmarks, or pricing are stated in the post.

  2. StepFunOfficialAI score20

    StepFun's Step 5 Preview targets finance tasks with FinStepBench evaluations

    AIStepFun says it is focusing Step 5 Preview on finance, judging it on verifying reliable information, reconciling conflicting reports, stating assumptions, and producing consistent, reproducible valuations. The post says the model is evaluated on FinStepBench, covering LiveSearch, CorporateValuation, and DeepResearch, and on FrontierFinance across six investment use cases.

    Image from @StepFun_ai's post
  3. StepFunOfficialAI score38

    StepFun previews Step 5 for large-scale research and analytical deliverables

    AIStepFun has previewed Step 5, an agent built for professional knowledge work spanning large-scale research, structured analysis, and interactive reporting. In one agent action, it coordinated 950 web fetches and assembled 300,000 monthly records across 1,000 locations over 25 years. In another, it produced a 17-sheet analytical workbook with source reconciliation, formulas, and trend models.

    Image from @StepFun_ai's post
  4. StepFunOfficialAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

    Why it matters: The post pairs a cost-versus-intelligence chart with specs and a later open-weights date, so readers can judge the cost tradeoff against named competitor models.

    Image from @StepFun_ai's post

Sep 18

Sep 18Fri
  1. VercelOfficialAI score22

    Jev Adopted Faster Than Any Model in AI Gateway History

    AIJev reached about 13% of teams on AI Gateway within its first day, which Vercel says is faster adoption than any other model in the gateway's history. That early uptake was roughly twice the GPT-5.6 family's and six times Fable 5.1's, according to the post.

    Image from @vercel's post
  2. WanOfficialAI score6

    Wan 3.0 raises output resolution by one level

    AIWan 3.0's output resolution has been raised by one level, according to a quoted post from @roco_kn_roco. The main post from Wan (@Alibaba_Wan) contains only emoji and adds no further details.

  3. Liquid AI · new models on Hugging FaceOfficialAI score55

    Liquid AI releases LFM2.5-VL-3B-DSpark drafter for faster vision-language decoding

    AILiquid AI released LFM2.5-VL-3B-DSpark, a speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The source reports decoding up to 2.66× faster on a single H100 with SGLang, up to 3.13× on Apple M5 Max with MLX-VLM, and up to 2.14× on Apple M3 Ultra with llama.cpp, with output unchanged under greedy decoding.

Sep 17

Sep 17Thu
  1. OpenBMBOfficialAI score36

    OpenBMB's MiniCPM5-2B runs offline on-device with 128K context

    AIOpenBMB's MiniCPM5-2B is a 2.5B-parameter model with native 128K context, offering hybrid Think and No-Think modes in one checkpoint. Users can download it from Hugging Face and run it fully offline on-device, as RunAnywhere demonstrated. In a demo, the model first called a puzzle impossible, then corrected itself and wrote a working verifier.

  2. xAI News (Grok)OfficialAI score42

    Grok Voice Transcribe 2.0 Doubles Accuracy of Predecessor at Same Price

    AIxAI released Grok Voice Transcribe 2.0, a speech-to-text model that is twice as accurate as Grok Voice Transcribe 1.0 at the same price, and ranks first for accuracy among 32 streaming models on the Artificial Analysis leaderboard. Batch transcription costs $0.10 per hour of audio and streaming $0.20 per hour, with diarization, timestamps, and key terms included. Existing Speech-to-Text API integrations gain the improvement with no code changes, and developers must pin grok-voice-transcribe-1.0 to stay on the older model during the transition.

  3. SenseTimeOfficialAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    AISenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

    Image from @SenseTime_AI's post
  4. OpenBMBOfficialAI score29

    Kahya-TTS: Turkish speech model fine-tuned from VoxCPM2 on 100 hours

    AIDeveloper Alican Kiraz fine-tuned OpenBMB's open-source VoxCPM2 voice model on nearly 100 hours of natural Turkish speech, creating Kahya-TTS for Turkish text-to-speech. The project shows how open-source voice models can be adapted to new languages and specialized datasets. The model is available on Hugging Face.

    Image from @OpenBMB's post
  5. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score46

    Ming-Image-0.1-Design-Layer splits flattened design images into RGBA layers

    AIinclusionAI has released Ming-Image-0.1-Design-Layer on Hugging Face, a model that decomposes a flattened design image into a requested number of RGBA layers using an image and a layer plan. The model runs at 1024 resolution (512 for faster processing) with 12 sampling steps, a CFG scale of 2.0, and BF16 precision on one CUDA GPU with 80 GiB VRAM. It is released under the MIT License.

  6. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score42

    inclusionAI releases Ming-Image-0.1-Design, a 6B text-to-image model for text-rich designs

    AIinclusionAI has released Ming-Image-0.1-Design, a 6B text-to-image model for UI, infographics, and posters that outputs RGBA images with transparent backgrounds. The model is available on Hugging Face and ModelScope under the MIT License. It runs at 2048 x 2048 with 12 sampling steps and a CFG scale of 1.0, validated on one CUDA GPU with 80 GiB VRAM.

  7. Z.aiOfficialAI score40

    GLM-5.3 helped build the inference stack serving GLM-5.3-Flash

    AIZ.ai reports that GLM-5.3 helped build and optimize the inference infrastructure for GLM-5.3-Flash. The system went from first successful run to production readiness in under two weeks, with end-to-end throughput tripling over the initial baseline. The team credited dense feedback from local correctness tests, execution traces, microbenchmarks, and end-to-end measurements for enabling targeted hypothesis testing.

Sep 16

Sep 16Wed
  1. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score55

    inclusionAI releases Realtime-Venus full-duplex audio-visual models on Hugging Face

    AIinclusionAI has published Realtime-Venus on Hugging Face with two 9B checkpoints: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for audio-only conversation. Both are built on MiniCPM-o 4.5 with a Qwen3-8B backbone and support full-duplex dialogue, proactive responses, and training-free long-video memory. The asynchronous Realtime-Venus-Harness runtime is hosted in a separate GitHub repository.

Sep 15

Sep 15Tue
  1. Tencent · new models on Hugging FaceOfficialAI score44

    Tencent releases WeVisDoc-4B, a document parser that leads OmniDocBench v1.6

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, converts page images into structured Markdown with LaTeX formulas and HTML tables. It scores 95.38 Overall on OmniDocBench v1.6 and a mean Overall of 75.54 across three PureDocBench tracks, ranking first among compared end-to-end parsers in all four reported settings. The model is available on Hugging Face and runs through vLLM, which requires version 0.11.1 or later.

  2. Tencent · new models on Hugging FaceOfficialAI score37

    Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models on Hugging Face

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, scores 95.38 Overall on OmniDocBench v1.6 and 75.54 mean Overall across three PureDocBench tracks. The end-to-end parser converts page images into structured Markdown with LaTeX formulas and HTML tables, and the 2B variant is also available. The repository provides vLLM serving scripts with a 32768-token default context and a Python client for batch processing.

  3. Chip HuyenXAI score40

    Jev model chooses from predefined outputs, promising very cheap inference

    AIChip Huyen praises an approach where models select from predefined values rather than generating freeform text, which she sees as useful for data labeling and fixed-action tasks. She notes that how reasoning would work is unclear, but the approach is very cheap because output tokens are free.

    Image from @chipro's post
  4. Google AI StudioOfficialAI score72

    Google releases Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps

    AIGoogle AI Studio released Gemini 3.8 Live, a native speech-to-speech model with an Extended Thinking variant, and made it available through the Live API. Gemini 3.5 Transcribe, released last month, supports 85+ languages with a reported 4.0% streaming and 2.6% non-streaming Word Error Rate, and accepts a custom vocabulary of up to 1,000 terms. Live API audio pricing is listed at $0.005/min for input and $0.018/min for output.

    Why it matters: The post lists concrete Live API capabilities, per-minute audio pricing, and transcription accuracy figures, helping developers weigh voice agent options against their own cascaded pipelines.

  5. koray kavukcuogluXAI score60

    Google Introduces Gemini 3.8 Live and Extended Thinking Voice Models

    AIGoogle announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, voice agents with reasoning capabilities. The post says the models take turns more seamlessly, think through complexity, and feel more natural to converse with.

    Why it matters: The post names two Gemini 3.8 Live variants and frames them around turn-taking and reasoning in voice conversation, the core change to weigh against earlier live models.

    Video from @koraykv's post
  6. LiveKitOfficialAI score31

    LiveKit Agents adds Gemini 3.8 Live for low-latency and extended-thinking voice agents

    AILiveKit Agents now supports Gemini 3.8 Live, letting developers choose gemini-3.8-live for low-latency audio or gemini-3.8-live-extended-thinking for longer asynchronous reasoning. Developers can switch between the two modes without changing their stack. The post points readers to LiveKit's documentation for more details.

    Video from @livekit's post
  7. Lewis Tunstall @ COLM 🌉XAI score30

    Periodic Labs advances toward cracking condensed matter physics superconductor problem

    AIPeriodic Labs, the team behind high-throughput materials labs in Menlo Park, reports progress on one of condensed matter physics' hardest problems. Its open-source model Neon, trained with mid-training and RL on 1,300 H200s plus months of lab data, surpasses GPT-6 Astra on the company's analysis benchmark. The work targets materials science challenges including superconductors, magnets, and semiconductors.

  8. Google AI StudioOfficialAI score46

    Google launches Gemini 3.8 Live and Extended Thinking dialogue models

    AIGoogle introduced two live dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, available through AI Studio and the Gemini API. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. The Extended Thinking variant targets high-complexity tasks with increased intelligence and multi-step reasoning.

    Video from @GoogleAIStudio's post
  9. Sundar PichaiXAI score42

    Google outlines AI for science, weather, languages, and economic research

    AIGoogle says it is focusing AI efforts on health, disaster and weather resilience, learning, and economic opportunity. Recent examples include AlphaGenome Atlas, which maps all 9B possible single-letter genetic changes across the human genome and is openly available to researchers, and WeatherNext 3, described as its most accurate and capable global weather AI model to date. The post also cites AI & Economy ATLAS, an open-access look at global AI usage, and says its translation services now cover nearly 300 languages spoken by 7B people.

    Image from @sundarpichai's post
  10. Logan KilpatrickXAI score44

    Google launches Gemini 3.8 Live audio models with 97-language support

    AIGoogle has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, described as new state-of-the-art live audio models available at frontier pricing and performance. The 3.8 Live model supports 97 languages with seamless switching between them, along with async tool calls.

    Image from @OfficialLoganK's post
  11. Google AIOfficialAI score72

    Google rolls out Gemini 3.8 Live and Extended Thinking across consumer, developer, and enterprise channels

    AIGoogle is rolling out Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking across several channels. Consumers get them in Search Live and Gemini Live, developers get public preview access through the Gemini API, and enterprises get private preview through Gemini Enterprise, with Customer Experience support coming soon.

    Why it matters: The post lays out where each Gemini 3.8 Live variant reaches consumers, developers, and enterprises, which clarifies access paths for a voice model release.

  12. Google AIOfficialAI score62

    Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking audio models

    AIGoogle AI announces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as its most advanced Gemini Audio models. Gemini 3.8 Live is built for scale, speed, and cost efficiency, handling mid-sentence interruptions, transitions across 97 languages, and visual context through Search Live. Gemini 3.8 Live Extended Thinking reasons and speaks in parallel, narrating its progress on multi-step tasks such as event planning.

    Why it matters: The post separates a low-cost real-time voice model from an extended-thinking variant, making the tradeoff between speed and reasoning depth clear to readers.

    Video from @GoogleAI's post
  13. Google DeepMindOfficialAI score72

    Google DeepMind releases Gemini 3.8 Live models for real-time voice agents

    AIGoogle DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents. Extended Thinking scores 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice, and 97.7% on Big Bench Audio. Gemini 3.8 Live is rolling out now in the Gemini API, Google AI Studio, and Search Live, with enterprise access in private preview.

    Why it matters: The release covers a voice model's benchmark results and availability across developer, enterprise, and consumer products, useful for judging voice agent options.

  14. Google AI StudioOfficialAI score72

    Google launches Gemini 3.8 Live and Extended Thinking voice models

    AIGoogle introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents that reason and speak simultaneously. The Extended Thinking version scores 82.6 on Artificial Analysis' Speech to Speech Quality Index and 97.7% on Big Bench Audio, while 3.8 Live targets scale and cost efficiency. Developers can access both through the Gemini API in Google AI Studio, and enterprise and consumer rollouts vary by product.

    Why it matters: The source names the two models, their access paths, and specific benchmark results, showing how the voice agent capabilities differ between the two tiers.

  15. Google DeepMindOfficialAI score33

    Gemini 3.8 Live Extended Thinking adds upgraded reasoning for real-time programming tutoring.

    AIGoogle DeepMind demonstrated 3.8 Live Extended Thinking acting as a programming tutor in Gemini Live. Both 3.8 Live models feature upgraded reasoning, near real-time visual understanding, automatic detection across 97 languages, and background tool calling that doesn't interrupt the chat. The Extended Thinking variant adds higher performance and precision for harder tasks and narrates its progress, and it is available in Gemini Live in the Gemini app or through the Gemini API via Google AI Studio.

    Video from @GoogleDeepMind's post