Really cool results from Sky Lab’s @Andy_ShuoYang to run massive LLMs on local GPUs!
Really cool results from Sky Lab’s @Andy_ShuoYang to run massive LLMs on local GPUs!
Really cool results from Sky Lab’s @Andy_ShuoYang to run massive LLMs on local GPUs!
We did a new Unsloth Desktop release! * Experimental auto compaction for any model * LAN & Remote Access tabs * Faster & smoother chatting * Over 200 PRs merged! For compaction, we use a mix of RAG + forced 1st turn RAG + tail - let us know how it goes! https://github.com/unslothai/unsloth/releases/tag/v0.1.801-beta
Google has released TIPS g/14 low-res (v1) on Hugging Face, a Text-Image Pre-training with Spatial awareness vision-language model with 1.1B vision parameters and 389M text parameters. The model produces spatially rich image features aligned with text embeddings at 224 resolution, under the Apache 2.0 license. It supports image encoding, text encoding, and zero-shot classification via the transformers library.
Google has released the original TIPS g/14 (v1) vision-language model on Hugging Face under Apache 2.0, with 1.1B vision parameters and 389M text parameters at 448 resolution. The TIPS family, presented at ICLR 2025, produces spatially rich image features aligned with text embeddings, and the release includes a low-res 224 variant.
We’re releasing new Qwen3.8-27B GGUFs with 10% higher accuracy. Unsloth Dynamic V3 outperforms others by >10% on Div-300, KLD & more benchmarks. We also release 1-bit quants that retain 77% accuracy. Run on 8GB RAM. Blog: https://unsloth.ai/docs/basics/dynamic-3.0-ggufs GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Google released google/tipsv1-so400m14, the original v1 So400m/14 checkpoint of TIPS, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 413M vision parameters and 448M text parameters at 448 resolution, and is licensed under Apache 2.0.
Google has published google/tipsv1-l14, the original v1 L/14 release of TIPS, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The L/14 variant has 304M vision parameters and 184M text parameters at 448 resolution, with an embedding dimension of 1024, and is licensed under Apache 2.0.
Google has published TIPS B/14 (v1) on Hugging Face, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 86M vision parameters and 110M text parameters at native 448 resolution, and is licensed under Apache 2.0. The release includes usage code for image and text encoding, zero-shot classification, and spatial feature visualization.
Build your own: https://github.com/nroadley/Creating-Oneshot-Hero-Landing-Pages
I find the "inception" pattern to be very useful in many agentic use cases. You can force the model to take an action when it thinks for too long by injecting a thought after a specified reasoning budget. Helps dealing with underspecified tasks which make the model reason for way too long.
The author built a Python agent that uses Gemini 3.7 Flash to control an Android emulator from raw screenshots, returning normalized 0–999 coordinates that are scaled to 1080x1920 pixels over ADB. In a test, the agent opened Chrome, closed popups, and solved one round of Wordle in two guesses without accessibility IDs or DOM access. The article presents the loop as usable for UI testing and task automation across native apps, webviews, and canvas interfaces, with code in an open-source quickstart repository.
Qwen3.8-27B GGUF has reached 1,000 likes in less than 24 hours! ❤️ It's now the #3 trending model on Hugging Face with 1M overall downloads. Run on 17GB RAM/VRAM setups via Unsloth! Model: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF GitHub: https://github.com/unslothai/unsloth
Cohere and Cohere Labs released North Small Translate 1.0 as open weights for research, a sparse Mixture-of-Experts model with 25B active and 218B total parameters. It is specialized for machine translation across 50 languages, with a 16K input and 16K output context. The chart shows a WMT26 all-languages score of 83.60, rising to 84.36 with the agentic multi-pass workflow, and the model is licensed CC BY-NC 4.0 with an acceptable use policy.
Why it matters: The model card lists the benchmark score, hardware needs, and license terms, which helps readers judge whether this translation model fits their use.
Z.ai says GLM-5.3 is its most capable model for cybersecurity tasks, with CyberGym at 84.5% versus 77.2% for GLM-5.2 and ExploitBench at 54.4% versus 24.4%. Access will begin with selected security partners in controlled settings, followed by broader access and API availability, with full open weights to be published after safety evaluations are complete. The company also launched the OpenVuln initiative to help open-source maintainers audit projects and coordinate disclosure.
Z.ai says GLM-5.3 is available now through GLM Coding Plan and ZCode. API access and open weights will follow in stages after rigorous safety evaluations.
MathForm-8B is an open-source autoformalization model from OpenBMB that translates natural-language mathematical statements into Lean 4. It was trained on FormalVerse through supervised fine-tuning, then reinforcement learning using Lean compilation and semantic-consistency feedback. The model is available on Hugging Face under Apache License 2.0 and can be served with Transformers, vLLM, or SGLang, using a recommended max_new_tokens of 16384.
DeepSeek has released DeepSeek Harness v0.1 in Developer Preview, opening the codebase under the MIT license for developers building agent harnesses. The harness is built on the Cordis meta-framework and treats models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI as plugins that can be mixed, matched, replaced, and extended.
Why it matters: The source specifies the MIT license and a plugin-based architecture covering models, tools, and sessions, which helps developers assess extensibility before adopting it.
Prime Intellect has released Prime Flash MoE, a set of Blackwell-optimized CUDA kernels for mixture-of-experts feed-forward layers. The kernels are up to 2.4× faster than PyTorch grouped GEMM and deliver about 2.3× speedup across the 4k–128k token range, and are integrated into its prime-rl framework. Two pipelines are offered: a fused single-kernel path for small problem sizes and a split three-kernel path for larger ones, supporting both bf16 and MXFP8.
Why it matters: The post explains how fusing MoE expert computation on Blackwell hardware avoids intermediate memory traffic, with benchmarks showing where fused and split pipelines each win.
MiniMax introduces Music 3.0, a music generation model that composes, arranges, performs, and produces a complete song from a creative concept and optional lyrics. The post describes an eight-layer RVQ tokenizer, a Hybrid-LM pairing an 8B Global LLM with a 0.6B Local LLM, and a flow-matching and Flow-VAE audio renderer. It says songs can run up to five minutes and that the model focuses on creative intent, arrangement, and vocal naturalness.
Why it matters: The post explains how the model's pipeline targets structure, acoustic detail, and vocal realism, which helps readers judge where open-weights music generation stands.
Meta announced it is opening the weights for Muse Glimmer, a 30B parameter dense model that can run locally. Muse Spark 1.2, described as its latest foundation model, will have its weights released soon. The author's interview with Mark Zuckerberg quotes him saying Llama 4 fell short of the trajectory he wanted and that the lab was rebuilt.
Nemotron 3.5 Lightning: Same architecture as 3.0 Nano, with added speculative decoding, and with the intelligence of 3.0 Super. ⚡️⚡️⚡️
Liquid AI released LFM2.5-2.6B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-2.6B target, on Hugging Face. In SGLang on a single H100 with batch size 1, mean decoding throughput rises from 323 to 864 tokens per second, about 2.67x, and on an Apple M4 Max via Metal it rises from 61 to 139 tokens per second, about 2.27x. Because the target verifies every proposed token, the output matches what LFM2.5-2.6B would generate alone.
Liquid AI released LFM2.5-1.2B-Instruct-DSpark, a 295.7M-parameter speculative-decoding draft model for the LFM2.5-1.2B-Instruct target on Hugging Face. On an H100 it averages 4.81 accepted tokens per step and runs about 2.10x faster across benchmarks, with about 2x speedup in SGLang and on-device Apple silicon support via Metal.
Thank you Mark, Alex and the whole Meta team for your contributions to open weight AI.
MiniMax Music 3 is a music generation model that creates complete songs up to five minutes long from lyrics and a music description. It pairs an 8B Global LLM for long-range structure with a 0.6B Local LLM for acoustic detail, outputting 32 kHz, 16-bit stereo WAV audio. The model is available on Hugging Face and supports SGLang-Omni, diffusers, and ComfyUI.
Shanghai AI Lab open-sourced Mobius, an architecture its authors compare to the RNN-to-Transformer shift in both token and knowledge dimensions. Against Transformers, the post claims about 4x faster reasoning, the same MMLU score with 40% less data, and 2x better compositional generalization. Mobius is supported by XTuner, LMDeploy, vLLM, and SGLang, and its experimental setup and training pipeline will be released later.
Alibaba's Qwen team has released Qwen3.8-27B on Hugging Face as a 27B dense model with native image and video understanding. The model card reports gains over Qwen3.6-27B on coding and agent benchmarks, including SWE-bench Pro at 61.7 versus 53.5. It adds reasoning_effort levels and preserve_thinking, and its hosted Qwen Cloud version is described as coming soon.
Why it matters: The model card gives per-benchmark comparisons with Qwen3.6-27B and named rivals, plus reasoning_effort and preserve_thinking controls for judging cost and agent behavior.
Prime Agent is a new open-source coding harness built on a persistent IPython kernel, a Recursive Language Model design, and Continual Harness state that the agent can create, read, update, and delete. Prime Intellect reports ARC-AGI-3 results of 95.5% RHAE Best@1 with Opus 5 and competitive long-context scores with the open-weights GLM-5.2 model.
Why it matters: The post explains how the RLM and Continual Harness designs let an agent write code against its own context, sub-agents, and harness state, with benchmark evidence.
Our challenge: Create a community-facing project for musicians using Stable Audio 3.0. Teams taking on our challenge can get started now. → GitHub: http://github.com/Stability-AI/stable-audio-3 → Model weights: http://huggingface.co/collections/stabilityai/stable-audio-3
Hot (summer) take: Europe should care, too, if we're ever going to compete. The only way to catch up in the age of AI is by a staunch commit to global, model agnostic open source collaboration, so that a breakthrough in Beirut is a breakthrough in Berlin. https://www.vox.com/politics/497534/chinese-ai-moonshot-kimi-deepseek-open-weight
Since releasing LongCat-2.0, we’ve kept building with it. Here are five creative projects, each made from a single prompt: 🧱 Voxel art 🌍 3D modeling 🌊 CG animation 🎞️ A landing page 🎮 A mini game Take a look 👇
Thinking Machines has released Inkling-Small, which the author says achieves performance comparable to Inkling at a quarter of its size. The model has 276B total parameters with 12B active, and its full weights are available, with fine-tuning on Tinker and chat in text, image, and audio on Tinker Playground.
Thinking Machines is releasing Inkling-Small, a model it says achieves performance comparable to Inkling at a quarter of its size. The model has 276B total parameters with 12B active, and full weights are available. It can be fine-tuned on Tinker or used for text, image, and audio chat in the Tinker Playground.
Poolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.
Alibaba NLP has released UEmbed-9B, a decoder-only multimodal embedding model built on Qwen3.5 9B that outputs both dense and SPLADE-style sparse embeddings from one forward pass. It supports text, image, video, and mixed-modal inputs for retrieval and multimodal search, and the family also includes 2B and 4B variants. The model is available on Hugging Face, with transformers and vLLM inference support.
Alibaba NLP has released UEmbed-4B, a decoder-only multimodal embedding model built on Qwen3.5 4B that outputs both dense and sparse embeddings from one forward pass. It handles text, image, video, and mixed-modal inputs for retrieval and visual-document search, and sparse activations map to vocabulary terms usable with inverted indexes. The model is available on Hugging Face in a family that also includes 2B and 9B variants.
Alibaba-NLP's UEmbed-2B, a decoder-only multimodal embedding model built on Qwen3.5 2B, produces both dense and SPLADE-style sparse embeddings from a single forward pass. It supports text, image, video, and mixed-modal inputs for retrieval, and the 4B and 9B variants are also available. The team reports state-of-the-art results on the text and agent tracks of MMEB-v3.
MiniMax released H3, an open-weights omni-modal model that generates video with native stereo audio up to 2K and 15 seconds. The system combines H3-Context-IR preprocessing, the H3-Base generator at 768p, and H3-Regenerate-2K for 2K output, with the Context-IR and 2K modules available only through API.
Why it matters: The source details a three-module pipeline and open weights with deployment paths, showing how a video model is served and reproduced locally.
Liquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, bidirectional encoders built on the LFM2 hybrid architecture and available on Hugging Face. They support an 8,192-token context and are designed for fine-tuning on classification and token-level tasks. On CPU, LFM2.5-Encoder-230M is the fastest model tested from 1K tokens up, running about 3.7x faster than ModernBERT-base at 8,192 tokens.
Kimi K3 is now available on Nebius Token Factory, which is named a Day 0 launch partner, through an OpenAI-compatible API and console. The quoted post says Artificial Analysis scores the open-weight model at 57 on its Intelligence Index, two points behind GPT-5.6 Sol (max), and lists up to 1M tokens of context.
Why it matters: The source names the cloud access route and an Artificial Analysis score of 57, letting readers compare Kimi K3 against GPT-5.6 Sol.