Skip to contentSkip to stories

Updated

Open source

Showing low-relevance items too. Hide low-relevance items

Sep 23

Sep 23Wed
  1. Julien ChaumondXAI score36

    Hugging Face releases JS package to browse LeRobot datasets directly

    AIHugging Face has released a new JavaScript package, huggingface/lerobot, that lets developers read LeRobot datasets on the Hub directly in the browser without downloading them. The post says coding agents can use it to build custom dataset viewers quickly, with an example UI implemented in about 600 lines of JS.

    Image from @julien_c's post
  2. ModelScopeOfficialAI score44

    NVIDIA releases Nemotron 3 Diarization for live speaker attribution

    AINVIDIA's Nemotron 3 Diarization is now available on ModelScope, labeling speakers and timestamps in streaming audio for up to eight speaker slots per conversation. The 99.2M-parameter model uses an end-to-end streaming architecture built on NVIDIA's Streaming Sortformer, running on Ampere, Hopper, and Blackwell GPUs via NeMo Speech C++. It is designed to pair with existing ASR systems such as Nemotron ASR, Parakeet, Canary, or Whisper to produce speaker-attributed transcripts.

    Image from @ModelScope2022's post
  3. QwenOfficialAI score60

    Qwen Intelligence launches three mobile agents and opens its benchmark suite

    AIAlibaba's Qwen launched Qwen Intelligence with three mobile agents: a Mobile Planner Agent, a Mobile-Use Agent, and a Mobile Creative Agent. The post reports benchmark results including MobileWorld 82.1, MobileWorld-Real 92.2, and AndroidDaily 97.2, plus a 90% end-to-end success rate, and says the MobilePA-Bench, MobileWorld, MobileWorld-Real, and MobileWorld-Safety benchmarks are open.

    Why it matters: The post names three mobile agents and their benchmark results, while also releasing the benchmark suite, so readers can check the claims against the reported figures.

    Image from @Alibaba_Qwen's post
  4. ModelScopeOfficialAI score40

    TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6

    AITeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model. It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.

    Image from @ModelScope2022's post
  5. ModelScopeOfficialAI score62

    Xiaomi MiMo-V2.6 open-sourced as a multimodal agent model family under MIT License

    AIXiaomi has released MiMo-V2.6 as an open model family under the MIT License, designed for large-scale reinforcement learning. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, with 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. The 1.02T-parameter MoE activates 42B parameters and supports text, image, video, and audio input with a 1M-token context.

    Why it matters: The post links benchmark results, parameter scale, and a multi-agent RL training run, giving readers concrete figures to compare against other open models.

    Image from @ModelScope2022's post
  6. ModelScopeOfficialAI score62

    Shanghai AI Lab and SJTU release open-weight 8.9B NCP-ArchPreview model under Apache 2.0

    AIShanghai AI Lab and SJTU's LUMIA Lab released NCP-ArchPreview, an 8.9B open-weight language model under Apache 2.0. The model reportedly reaches OLMo-3-7B's final Stage 1 loss using 51.3% of the tokens from the 5.73T Dolma 3 corpus, a 1.95× convergence gain. Its concept module jointly predicts tokens and concepts, and domain adaptation updates only its 17M parameters while the token backbone stays frozen.

    Why it matters: The post pairs an Apache 2.0 open-weight release with training-efficiency figures, showing how the concept module adapts to new domains with few trainable parameters.

    Image from @ModelScope2022's post

Sep 22

Sep 22Tue
  1. ModelScopeOfficialAI score62

    inclusionAI open-sources Ming-Image-0.1-Design models for visual design

    AIinclusionAI open-sources the Ming-Image-0.1-Design family, two complementary 6B models for visual-design workflows, under an MIT License. Design generates complete UIs, dashboards, infographics, and posters up to 2048×2048 with native transparent RGBA output, and Layer decomposes flattened graphics into independently editable RGBA layers.

    Why it matters: The post separates a design-generation model from a layer-decomposition model, letting users compare two distinct visual-design workflows under one MIT license.

    Image from @ModelScope2022's post
  2. Daniel HanXAI score22

    Unsloth Desktop hotfix adds Qwen-Image-2.1 image editing and fixes

    AIUnsloth Desktop received a hotfix update adding image editing for Qwen-Image-2.1. The update also fixes diffusers update issues, GGUF loading failures for Qwen-Image, and black artifacts during diffusion on A100 and consumer GPUs. Users should receive a banner prompting them to update.

  3. Unsloth AIOfficialAI score26

    Qwen-Image-2.1 FP8 and GGUF quants now run in Unsloth Desktop

    AIUnsloth announced that Qwen-Image-2.1 FP8 and GGUF quantized versions should now run properly in Unsloth Desktop. The app supports both image generation and image editing with these quants. Further details are available on the Unsloth GitHub repository.

    Image from @UnslothAI's post
  4. Fireworks AI BlogOfficialAI score46

    Fireworks ARCv3 cuts RL weight-update payloads nearly 50% for cross-region training

    AIFireworks released ARCv3, a lossless compressor for BF16 weight-update deltas sent from trainers to RL rollout machines. Across 1,000 production RL deltas, ARCv3 produced payloads nearly 50% smaller than ARCv2, averaging about 0.19% of the BF16 weight size versus 0.36%. ARCv3 is available through the Fireworks Training API as fireworks-delta-compression.

  5. Boris ChernyXAI score42

    Boris Cherny Uses Opus 5.5 to Formally Verify Claude Agent SDK

    AIBoris Cherny used Opus 5.5 to formally verify the Claude Agent SDK with Lean, and short prompts produced 16 PRs fixing bugs and race conditions. He also combines Lean and TLA+ to find issues in data flow, concurrency, and state management, and says Claude is strong in both languages even though he does not know them well.

    Video from @bcherny's post
  6. whXAI score34

    MiMo-V2.6 paper details data and RL results for open model

    AIThe MiMo-V2.6 paper thread reports on the newest open model, which also streams its RL run, focusing on data and RL experimental results rather than architecture. The Pro model reportedly rose from 58.41 to 72.57 on DeepSWE after RL, with the top published DeepSWE score cited at 74.

    Image from @nrehiew_'s post
  7. Google GemmaOfficialAI score31

    Deploy DiffusionGemma-Jev on Google Cloud Run with one command

    AIGoogle Gemma says DiffusionGemma-Jev (djev) can now be deployed as a Jev API-compatible endpoint on Google Cloud Run with a single command. The post reports about 35-60 ms single-step latency and roughly 100-123 requests/sec at batch size 32, at about $3/hr that drops to $0 when idle.

    Video from @googlegemma's post
  8. Google GemmaOfficialAI score22

    Google Gemma credits DiffusionGemma-Jev deployment on Cloud Run

    AIGoogle Gemma credits @mmastrac and @dylayed for work on DiffusionGemma-Jev (djev), a Jev API-compatible endpoint. Per @dylayed, djev can be deployed to Google Cloud Run with a single gcloud command, at roughly $3/hr while active and $0 when idle.

  9. Comfy BlogOfficialAI score42

    ComfyUI Speeds Up MiniMax H3 Video VAE Encoding and Decoding

    AIComfyUI's update makes the MiniMax H3 video VAE encode up to about 2.2x faster and decode 1.4-2.7x faster, cutting a 1344x768, 129-frame round trip on an RTX 5090 from 24.3 to 12.7 seconds. The gains come from a fused encoder kernel enabled by default, fp16 accumulation support in a custom convolution, and an int8 decoder, and the source says the changes are visually lossless to the eye. Users need ComfyUI v0.36.0 or above, and the int8 VAE file is a drop-in replacement for the standard one.

  10. StepFunOfficialAI score43

    StepFun open-sources onPanda for token-level LLM annotation and inspection

    AIStepFun has open-sourced onPanda, a tool used internally for LLM data annotation and model inspection, letting users correct tokens and let models continue. The company reports a 52% lower median annotation time versus manual post-editing, with SFT and preference data combined in one workflow. It also supports token probability and top-k inspection, token-by-token decoding control, and browser-based testing across SVG generation, web development, and agent tasks.

  11. LlamaIndex 🦙OfficialAI score22

    LiteParse v2.14.6 parses text PDFs about 25% faster locally

    AILlamaIndex released LiteParse v2.14.6, an open-source PDF-to-Markdown parser that processes text-based PDFs about 25% faster. On realistic documents it handled pages at 2.8ms per page, 1.5 times faster than the next-fastest local parser. It runs locally in Python, Node.js, Rust, or directly in the browser.

    Image from @llama_index's post
  12. StepFunOfficialAI score31

    Step Code installs on macOS, Linux, and WSL via one command

    AIStepFun released an install script for Step Code that runs on macOS, Linux, or WSL with a single curl command. WSL is the recommended route on Windows, while PowerShell support is currently in beta. The project accepts issues and pull requests on GitHub.

  13. StepFunOfficialAI score52

    StepFun releases Step Code v0.1.0 as an open-source coding CLI

    AIStepFun has released Step Code v0.1.0, an open-source command-line tool under the MIT License that covers reading and editing code, running tests, and shipping from one CLI. The post reports 80.9% on Terminal-Bench 2.1 and 73.3% on Multi-Frame, a 150-task long-horizon benchmark from StepFun. It also includes one-command static site publishing with StepPage and links the GitHub repository.

    Image from @StepFun_ai's post
  14. Daniel HanXAI score42

    Qwen-Image-2.1 runs locally in Unsloth Desktop via INT8, FP8, GGUF

    AIDaniel Han says Qwen-Image-2.1 works in Unsloth Desktop through INT8, FP8, and GGUF builds, with Unsloth also releasing dynamic GGUFs for it. Pinned RAM offloading lets INT8 and FP8 fit under 6–8GB of VRAM while remaining relatively fast. The linked Unsloth post says the 7B model runs on 12GB VRAM and performs on par with Nano Banana 2.0.

  15. Unsloth AIOfficialAI score70

    Qwen-Image-2.1 runs locally on 12GB VRAM using Unsloth GGUFs

    AIUnsloth says the 7B Qwen-Image-2.1 text-to-image and editing model can run locally on 12GB VRAM using its GGUF builds. It also states that the model performs on par with Nano Banana 2.0, and that Dynamic FP8 can run on 6GB of VRAM via offloading for higher quality. The image lists int8 at 7.26 GB with mean LPIPS 0.064 and fp8 at 7.12 GB with mean LPIPS 0.112, and says int8 is the default.

    Why it matters: The post gives concrete local-run settings, VRAM figures, and GGUF and FP8 options, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  16. Sebastian RaschkaXAI score62

    Xiaomi MiMo-V2.6-Pro tops open-weight benchmarks with simple attention design

    AIXiaomi's MiMo-V2.6-Pro ranks first among open-weight models on the Artificial Analysis Intelligence Index with a score of 46. The author attributes its standing mainly to a training data and post-training recipe that increased agent tasks and used an agentic grader for rewards, rather than its plain Grouped Query Attention and Sliding Window Attention design with a 128-token window.

    Image from @rasbt's post
  17. OpenBMBOfficialAI score59

    VoxWeft runs real-time interpretation locally on Apple Silicon using VoxCPM2

    AIOpenBMB highlights VoxWeft, an open-source simultaneous interpretation system for Apple Silicon built by developer @HenryZ30734018 on an MLX implementation of VoxCPM2. The system turns live speech into translated speech on-device, with first audio streaming in about 170 ms on an M5 MacBook. VoxCPM2 generates speech in 30 languages, supports direct language-pair interpretation, and clones a target voice from about 5 seconds of reference audio.

    Video from @OpenBMB's post
  18. Black Forest Labs · new models on Hugging FaceOfficialAI score62

    Black Forest Labs releases FLUX 3 Action, a 7B open-weights robot world action model

    AIBlack Forest Labs released FLUX 3 Action, an open-weights 7B world action model that outputs robot joint commands from camera frames, robot state, and a text instruction. On the RoboLab-120 benchmark it reports 42.92% task success, ahead of Cosmos3-Nano-Policy at 36.8% and π0.5 at 28.0%. The model is fine-tuned on DROID, is distributed under the FLUX Kommunity License v.1.0, and runs in about 32 GB of GPU memory in bfloat16.

    Why it matters: The model card gives a benchmark comparison, parameter counts, and an action contract, so readers can judge how it compares with existing robot policies.

  19. Black Forest Labs · new models on Hugging FaceOfficialAI score58

    Black Forest Labs releases open-weights FLUX 3 Action SO-101 robot policy

    AIBlack Forest Labs has published FLUX 3 Action SO-101 on Hugging Face as an open-weights 7B world action model. It takes two camera frames, the robot state, and a text instruction, then returns the next 42 actions with predicted video frames, with 32 executed at 30 Hz before replanning. The card also provides a rank-32 LoRA fine-tuning recipe for user datasets and states that the application must enforce joint velocity, force, and workspace limits.

  20. Black Forest Labs · new models on Hugging FaceOfficialAI score60

    Black Forest Labs releases FLUX 3 Action base weights for robot adaptation

    AIBlack Forest Labs has released flux-3-action-base, an open-weights 7B world action model that takes camera frames, robot state, and a text instruction to output the next action chunk. The release is an adaptation component rather than a complete robot policy, and new embodiments require their own action heads. The source says the weights are paired with shared video VAE and Qwen3-VL-4B-Instruct text encoders and is governed by the FLUX Kommunity License v.1.0.

    Why it matters: The source separates the adaptation base from full robot policies and states the shared encoders and new-embodiment requirements, which clarifies what developers must still build for their robots.

  21. OpenBMBOfficialAI score20

    OpenBMB praises MiniCPM5-2B workers in multi-agent invoice reconciliation

    AIOpenBMB thanked a developer for testing MiniCPM5-2B as a worker in a multi-agent workflow handling invoice matching, short payments, duplicate references, and disputes through tool calls. The background post says GPT-6 Astra coordinated the MiniCPM5-2B workers, verifying 32 synthetic invoices in 67.8 seconds with 232 executed tool calls. The demo does not move money.

  22. X.PINXAI score46

    Moonshot's Kimi K3 now available on Amazon Bedrock

    AIMoonshot's Kimi K3 is now available on Amazon Bedrock, with its license requiring a paid agreement for model-hosting businesses and affiliates above $20M in annual revenue. AWS says customer data stays within its cloud, is not shared with Moonshot or used for training, and inference requests have zero data retention. Neither company disclosed financial terms.

    Image from @thexpin's post

Sep 21

Sep 21Mon
  1. StepFunOfficialAI score58

    StepFun's Step 5 Preview scores 44 on Intelligence Index at lower cost

    AIStepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index at about $0.72 per task, matching Kimi K3 (max) at roughly 2.8x lower cost. The source reports strong reasoning results, including 46% on Humanity's Last Exam, but places it behind Qwen3.8 Max and GLM-5.3 (max) on agentic evaluations. Open weights are planned for October 15.

  2. Tencent HyOfficialAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.