Skip to contentSkip to stories

Updated

#Open source/Repo

Showing low-relevance items too. Hide low-relevance items

Sep 23

Sep 23Wed
  1. QwenOfficialAI score60

    Qwen Intelligence launches three mobile agents and opens its benchmark suite

    AIAlibaba's Qwen launched Qwen Intelligence with three mobile agents: a Mobile Planner Agent, a Mobile-Use Agent, and a Mobile Creative Agent. The post reports benchmark results including MobileWorld 82.1, MobileWorld-Real 92.2, and AndroidDaily 97.2, plus a 90% end-to-end success rate, and says the MobilePA-Bench, MobileWorld, MobileWorld-Real, and MobileWorld-Safety benchmarks are open.

    Image from @Alibaba_Qwen's post
  2. ModelScopeOfficialAI score62

    Xiaomi MiMo-V2.6 open-sourced as a multimodal agent model family under MIT License

    AIXiaomi has released MiMo-V2.6 as an open model family under the MIT License, designed for large-scale reinforcement learning. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, with 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. The 1.02T-parameter MoE activates 42B parameters and supports text, image, video, and audio input with a 1M-token context.

    Image from @ModelScope2022's post

Sep 22

Sep 22Tue
  1. ModelScopeOfficialAI score62

    inclusionAI open-sources Ming-Image-0.1-Design models for visual design

    AIinclusionAI open-sources the Ming-Image-0.1-Design family, two complementary 6B models for visual-design workflows, under an MIT License. Design generates complete UIs, dashboards, infographics, and posters up to 2048×2048 with native transparent RGBA output, and Layer decomposes flattened graphics into independently editable RGBA layers.

    Image from @ModelScope2022's post
  2. Unsloth AIOfficialAI score26

    Qwen-Image-2.1 FP8 and GGUF quants now run in Unsloth Desktop

    AIUnsloth announced that Qwen-Image-2.1 FP8 and GGUF quantized versions should now run properly in Unsloth Desktop. The app supports both image generation and image editing with these quants. Further details are available on the Unsloth GitHub repository.

    Image from @UnslothAI's post
  3. Google GemmaOfficialAI score31

    Deploy DiffusionGemma-Jev on Google Cloud Run with one command

    AIGoogle Gemma says DiffusionGemma-Jev (djev) can now be deployed as a Jev API-compatible endpoint on Google Cloud Run with a single command. The post reports about 35-60 ms single-step latency and roughly 100-123 requests/sec at batch size 32, at about $3/hr that drops to $0 when idle.

    Video from @googlegemma's post
  4. Google GemmaOfficialAI score22

    Google Gemma credits DiffusionGemma-Jev deployment on Cloud Run

    AIGoogle Gemma credits @mmastrac and @dylayed for work on DiffusionGemma-Jev (djev), a Jev API-compatible endpoint. Per @dylayed, djev can be deployed to Google Cloud Run with a single gcloud command, at roughly $3/hr while active and $0 when idle.

  5. StepFunOfficialAI score43

    StepFun open-sources onPanda for token-level LLM annotation and inspection

    AIStepFun has open-sourced onPanda, a tool used internally for LLM data annotation and model inspection, letting users correct tokens and let models continue. The company reports a 52% lower median annotation time versus manual post-editing, with SFT and preference data combined in one workflow. It also supports token probability and top-k inspection, token-by-token decoding control, and browser-based testing across SVG generation, web development, and agent tasks.

  6. LlamaIndex 🦙OfficialAI score22

    LiteParse v2.14.6 parses text PDFs about 25% faster locally

    AILlamaIndex released LiteParse v2.14.6, an open-source PDF-to-Markdown parser that processes text-based PDFs about 25% faster. On realistic documents it handled pages at 2.8ms per page, 1.5 times faster than the next-fastest local parser. It runs locally in Python, Node.js, Rust, or directly in the browser.

    Image from @llama_index's post
  7. StepFunOfficialAI score31

    Step Code installs on macOS, Linux, and WSL via one command

    AIStepFun released an install script for Step Code that runs on macOS, Linux, or WSL with a single curl command. WSL is the recommended route on Windows, while PowerShell support is currently in beta. The project accepts issues and pull requests on GitHub.

  8. StepFunOfficialAI score52

    StepFun releases Step Code v0.1.0 as an open-source coding CLI

    AIStepFun has released Step Code v0.1.0, an open-source command-line tool under the MIT License that covers reading and editing code, running tests, and shipping from one CLI. The post reports 80.9% on Terminal-Bench 2.1 and 73.3% on Multi-Frame, a 150-task long-horizon benchmark from StepFun. It also includes one-command static site publishing with StepPage and links the GitHub repository.

    Image from @StepFun_ai's post
  9. OpenBMBOfficialAI score59

    VoxWeft runs real-time interpretation locally on Apple Silicon using VoxCPM2

    AIOpenBMB highlights VoxWeft, an open-source simultaneous interpretation system for Apple Silicon built by developer @HenryZ30734018 on an MLX implementation of VoxCPM2. The system turns live speech into translated speech on-device, with first audio streaming in about 170 ms on an M5 MacBook. VoxCPM2 generates speech in 30 languages, supports direct language-pair interpretation, and clones a target voice from about 5 seconds of reference audio.

    Video from @OpenBMB's post
  10. Black Forest Labs · new models on Hugging FaceOfficialAI score62

    Black Forest Labs releases FLUX 3 Action, a 7B open-weights robot world action model

    AIBlack Forest Labs released FLUX 3 Action, an open-weights 7B world action model that outputs robot joint commands from camera frames, robot state, and a text instruction. On the RoboLab-120 benchmark it reports 42.92% task success, ahead of Cosmos3-Nano-Policy at 36.8% and π0.5 at 28.0%. The model is fine-tuned on DROID, is distributed under the FLUX Kommunity License v.1.0, and runs in about 32 GB of GPU memory in bfloat16.

    Why it matters: The model card gives a benchmark comparison, parameter counts, and an action contract, so readers can judge how it compares with existing robot policies.

  11. Black Forest Labs · new models on Hugging FaceOfficialAI score58

    Black Forest Labs releases open-weights FLUX 3 Action SO-101 robot policy

    AIBlack Forest Labs has published FLUX 3 Action SO-101 on Hugging Face as an open-weights 7B world action model. It takes two camera frames, the robot state, and a text instruction, then returns the next 42 actions with predicted video frames, with 32 executed at 30 Hz before replanning. The card also provides a rank-32 LoRA fine-tuning recipe for user datasets and states that the application must enforce joint velocity, force, and workspace limits.

Sep 21

Sep 21Mon
  1. Xiaomi MiMoOfficialAI score31

    Xiaomi's MiMo-V2.6-Pro reaches top 10 on Code Arena WebDev

    AIArena says Xiaomi's MiMo-V2.6-Pro debuted at about #10 overall on Code Arena: WebDev with a 1628-point AutoEval score, tying Claude Fable 5 (High). That is a 153-point gain over MiMo-V2.5-Pro's 1475, and it ranks about #3 among open-weights models under an MIT license. Arena notes the score is early, based on a reward model rather than live human votes, so rankings may shift as more votes arrive.

  2. Xiaomi MiMoOfficialAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    AIXiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    Why it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

    Image from @XiaomiMiMo's post
  3. Apple · new models on Hugging FaceOfficialAI score46

    Apple releases LensVLM-9B, a vision-language model for compressed text images

    AIApple has released LensVLM-9B on Hugging Face, a 9B-parameter Vision Language Model that scans compressed images of text and selectively expands relevant pages to their uncompressed form. The repository provides a demo script and supports compression settings of 5x, 10x, and 15x. Model files are under the Apple Machine Learning Research Model License, and the accompanying source code is distributed separately under the Apple Sample Code License.

  4. Engineering at MetaOfficialAI score39

    Meta Open-Sources Rebalancer, a Library for Solving Assignment Problems

    AIMeta has open-sourced Rebalancer, an assignment-problem solver it has used for over nine years to allocate resources across its infrastructure. The library separates problem specification, in-memory storage, solving, and debugging, and translates problems into expression graphs solved via local search or mixed integer programs using FICO Xpress, Gurobi, or the open-source HiGHS solver.

  5. Xiaomi MiMo · new models on Hugging FaceOfficialAI score67

    Xiaomi releases MiMo-V2.6-Flash-RL, a 309B sparse MoE model with 1M context

    AIXiaomi released MiMo-V2.6-Flash-RL, an efficiency-balanced checkpoint in its MiMo-V2.6 series, on Hugging Face. The model is a sparse MoE with 309B total and 15B activated parameters, supports text, image, video, and audio input, and offers a 1M-token context. The technical report says it was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs its benchmark tables with the RL training method, which helps readers judge how the checkpoint's scores relate to its training approach.

  6. Microsoft ResearchOfficialAI score50

    Microsoft Research open-sources RetroChimera, a retrosynthesis model published in Nature

    AIMicrosoft Research published RetroChimera, a retrosynthesis framework that combines the R-SMILES 2 Transformer model and the NeuralLoc graph neural network through learned ensembling to propose synthesis routes for small molecules. In blind tests, PhD-level chemists preferred its individual reaction predictions over those from preceding models and recorded literature reactions. The implementation and weights are open-sourced for researchers developing new medicinal molecules and materials.

  7. OpenBMBOfficialAI score23

    Developer builds local MiniCPM News Desk for traceable AI news briefings

    AIDeveloper Mark Fenner built MiniCPM News Desk, a local-first news briefing system powered by MiniCPM5-2B. The model selects key passages from official AI and technology sources, and a rule-based editorial layer preserves dates and context before assembling a daily recap. Invalid or incomplete outputs are rejected, and the full pipeline runs locally without a hosted-model fallback.

    Image from @OpenBMB's post
  8. MiniMax Design (H3)OfficialAI score22

    Community speeds up MiniMax H3 video generation with sparse attention

    AIA community developer integrated the Jev method into MiniMax H3 to sparsify attention, deciding per layer which parts to keep. On an RTX 4070, video generation time dropped from 6 min 7 sec to 3 min 34 sec, a 41.7% reduction. The post notes that Jev selected sparsity rates of 1%, 3%, 5%, and 10% across 49 layers in 4-step generation.

Sep 20

Sep 20Sun
  1. swyxXAI score22

    Jev Podcast Episode Announced by Latent Space Host swyx

    AIswyx announced a Latent Space podcast episode featuring Jev, subscribable on Apple and YouTube, and thanked guests Allen Park and Ke. A quoted post from @CompleteSkeptic claims Jev is a frontier model with 20-200x faster speed and 40-400x lower cost, but this post itself adds no verified details.

    Image from @swyx's post
  2. LMSYS OrgOfficialAI score32

    RLinf adds Cosmos3 support with SGLang, boosting evaluation throughput 3.33x

    AIRLinf, an open-source framework for embodied intelligence and AI agents, now supports Cosmos3 from fine-tuning through robot evaluation. With SGLang inference, it delivers 3.33x end-to-end evaluation throughput, batching inference for 128 parallel environments on 8 GPUs across 500 episodes of the full LIBERO-10 evaluation. RLinf also overlaps CPU simulation with GPU inference to reduce waiting between stages.

    Image from @lmsysorg's post
  3. OpenBMBOfficialAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post
  4. Qwen · new models on Hugging FaceOfficialAI score62

    Qwen releases Qwen-Image-2.1 prompt rewriter for image editing on Hugging Face

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B visual generation parameters. The Hugging Face page for Qwen-Image-2.1-PE-I2I is a fine-tuned Qwen3.5-VL 9B prompt rewriter that turns vague editing instructions and input images into precise editing prompts, supporting up to 10 reference images.

    Why it matters: The model card documents usage with transformers and diffusers, letting readers see how the editing prompt rewriter connects to the generation pipeline.

  5. Qwen · new models on Hugging FaceOfficialAI score62

    Qwen releases open-source Qwen-Image-2.1 with a prompt rewriting model

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with a 7B-parameter visual generation component. The release also includes Qwen-Image-2.1-PE-T2I, a fine-tuned Qwen3.5-VL 9B model that rewrites brief image requests in any language into detailed English prompts with a recommended aspect ratio.

    Why it matters: The release pairs a 7B visual generation component with a separate prompt rewriting model, showing how a brief image request becomes a detailed English prompt before rendering.

Sep 19

Sep 19Sat
  1. OpenBMBOfficialAI score34

    OpenBMB's 2B MiniCPM5 powers a local personal news desk

    AIOpenBMB's 2B-parameter MiniCPM5 model runs as a local news desk on an older i5-9400F PC with 16GB RAM and no cloud API. The developer built a system that collects official sources hourly and sends a 24-hour Telegram recap with a lead story and links.

Sep 18

Sep 18Fri
  1. LM StudioOfficialAI score22

    Splash engine released as open source on GitHub

    AIThe Splash engine, posted by LM Studio, is now available as open source on GitHub. The post provides only a link to the incoai/splash repository and includes no further technical details.

Sep 17

Sep 17Thu
  1. OpenBMBOfficialAI score36

    OpenBMB's MiniCPM5-2B runs offline on-device with 128K context

    AIOpenBMB's MiniCPM5-2B is a 2.5B-parameter model with native 128K context, offering hybrid Think and No-Think modes in one checkpoint. Users can download it from Hugging Face and run it fully offline on-device, as RunAnywhere demonstrated. In a demo, the model first called a puzzle impossible, then corrected itself and wrote a working verifier.

  2. ChatGPTOfficialAI score38

    ChatGPT now available as an add-in inside Microsoft Word

    AIOpenAI's ChatGPT can now be added to Microsoft Word, letting users turn rough notes into first drafts, rework tangled paragraphs, proofread, and get suggested edits without leaving their document. The post also says it can flag formatting issues.

    Video from @ChatGPT's post