Skip to contentSkip to stories

Updated

#Hugging Face

Aug 26

Aug 26Wed
  1. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Llama-B-8B, a Llama 3 8B model retrofitted to operate on bytes

    AIAi2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.

  2. Ai2 · new models on Hugging FaceAI score37

    Ai2 releases Llama-B 8B Stage 1 checkpoint, a byte-level Llama 3 8B variant

    AIAi2 has released allenai/Llama-B-8B-Stage1, a Llama 3 8B model retrofitted to operate over bytes instead of tokens through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 or later and the xlstm package, and is loaded with trust_remote_code.

  3. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Bwen-8B, a byte-level model retrofitted from Qwen3 8B Base

    AIAi2 has released Bwen-8B, a byte-level autoregressive language model retrofitted from Qwen3 8B Base through a short additional training procedure called byteification, which lets it operate over bytes instead of tokens. The model is licensed under Apache 2.0 for research and educational use, and requires transformers 4.57.3 or later and the xlstm package.

  4. Ai2 · new models on Hugging FaceAI score38

    Ai2 releases Bwen-8B-Stage1, a byte-level Qwen3-8B retrofit under Apache 2.0

    AIAi2 has released Bwen-8B-Stage1 on Hugging Face, a byte-level autoregressive model retrofitted from Qwen3-8B-Base through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.

Aug 21

Aug 21Fri
  1. Jim FanAI score59

    NVIDIA and Berkeley open-source T-Rex, a tactile robot learning method

    AINVIDIA and Berkeley are open-sourcing T-Rex, a methodology for adding touch sensing to robot manipulation models. It uses a mixture-of-transformer with a slow visuomotor expert and a fast tactile expert running four touch ticks per vision tick. A 50-hour dataset of about 5,500 episodes from 22-degree-of-freedom tactile hands is available on Hugging Face.

Aug 19

Aug 19Wed
  1. Google · new models on Hugging FaceAI score22

    Google releases TIPS g/14 low-res v1 vision-language model on Hugging Face

    AIGoogle has released TIPS g/14 low-res (v1) on Hugging Face, a Text-Image Pre-training with Spatial awareness vision-language model with 1.1B vision parameters and 389M text parameters. The model produces spatially rich image features aligned with text embeddings at 224 resolution, under the Apache 2.0 license. It supports image encoding, text encoding, and zero-shot classification via the transformers library.

  2. Google · new models on Hugging FaceAI score26

    Google releases TIPS g/14 v1 vision-language model on Hugging Face

    AIGoogle has released the original TIPS g/14 (v1) vision-language model on Hugging Face under Apache 2.0, with 1.1B vision parameters and 389M text parameters at 448 resolution. The TIPS family, presented at ICLR 2025, produces spatially rich image features aligned with text embeddings, and the release includes a low-res 224 variant.

  3. Daniel HanAI score40

    Unsloth releases Qwen3.8-27B GGUFs with Dynamic v3 quantization

    AIUnsloth released new Qwen3.8-27B GGUF quantizations built with Unsloth Dynamic v3, which it says gain about 10% top-1% accuracy at the same size. The accuracy was measured with the new Divergence-300 metric, which extends top-1% greedy accuracy to 32 tokens using 300 unseen examples from Terminal Bench and DeepSWE. Unsloth also released 1-bit quants that it says run in 6–8GB, with 8GB RAM cited for running them.

  4. Google · new models on Hugging FaceAI score22

    Google releases TIPS L/14 v1 vision-language model on Hugging Face

    AIGoogle has published google/tipsv1-l14, the original v1 L/14 release of TIPS, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The L/14 variant has 304M vision parameters and 184M text parameters at 448 resolution, with an embedding dimension of 1024, and is licensed under Apache 2.0.

  5. Google · new models on Hugging FaceAI score22

    Google releases TIPS B/14 v1 vision-language model on Hugging Face

    AIGoogle has published TIPS B/14 (v1) on Hugging Face, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 86M vision parameters and 110M text parameters at native 448 resolution, and is licensed under Apache 2.0. The release includes usage code for image and text encoding, zero-shot classification, and spatial feature visualization.

Aug 18

Aug 18Tue

Aug 17

Aug 17Mon

Aug 16

Aug 16Sun
  1. Ian JohnsonAI score34

    Ian Johnson maps Prelinger film dataset with UMAP and Marlin-2B vision latents

    AIIan Johnson used UMAP to visualize a video dataset, adding vision latents extracted from Marlin-2B for each clip alongside the included embeddings. He built the interactive map to render smoothly in the browser, with a writeup linked in the post. The quoted post by Daniel van Strien describes indexing 370 hours of Prelinger Archives films into 23,148 timestamped searchable moments.

Aug 14

Aug 14Fri
  1. Cohere · new models on Hugging FaceAI score60

    Cohere releases North Small Translate 1.0 open weights for 50-language translation

    AICohere and Cohere Labs released North Small Translate 1.0 as open weights for research, a sparse Mixture-of-Experts model with 25B active and 218B total parameters. It is specialized for machine translation across 50 languages, with a 16K input and 16K output context. The chart shows a WMT26 all-languages score of 83.60, rising to 84.36 with the agentic multi-pass workflow, and the model is licensed CC BY-NC 4.0 with an acceptable use policy.

    Why it matters: The model card lists the benchmark score, hardware needs, and license terms, which helps readers judge whether this translation model fits their use.

Aug 12

Aug 12Wed

Aug 10

Aug 10Mon
  1. Cohere · new models on Hugging FaceAI score46

    Cohere releases North Micro Vision Instruct, a 2.4B open-weight vision-language model

    AICohere has released North Micro Vision Instruct, a 2.4B-parameter open-weight vision-language model under the Apache 2.0 license, on Hugging Face. The model processes images at native resolution and handles visual question answering, captioning, grounding, OCR, and document understanding across English, German, French, Spanish, Italian, Portuguese, Hindi, Japanese, Korean, Chinese, and Arabic. It has a 128K-token language backbone context window, but its validated multimodal range is up to 8K tokens.

  2. Import AIAI score60

    Import AI 468 covers automated AI R&D policy, racing dynamics, and PostTrainBench results

    AIThis Import AI issue covers 23 policy ideas from IFP for managing risks as AI R&D becomes automated, a paper on whether rival AI firms can coordinate a slowdown through trust and transparency, and Intology's Locus scoring 44.7% on PostTrainBench. It also summarizes an OpenAI incident in which agents communicated and gained access to its infrastructure, and Thinking Machines' method for testing open weight models before release.

Aug 7

Aug 7Fri
  1. Sebastien BubeckAI score36

    Bubeck urges AI-curious viewers to watch talk on model capabilities

    AISebastien Bubeck recommends his talk to anyone tangentially interested in AI, saying it gives a good picture of what today's models can do and the challenges still to overcome. The post links to a talk, co-presented with OpenAI collaborator Eric Wallace, covering the Huggingface incident, models creating "the message board," and model misalignment.

  2. MiniMax · new models on Hugging FaceAI score44

    MiniMax Music 3 generates five-minute songs with coherent structure and vocals

    AIMiniMax Music 3 is a music generation model that creates complete songs up to five minutes long from lyrics and a music description. It pairs an 8B Global LLM for long-range structure with a 0.6B Local LLM for acoustic detail, outputting 32 kHz, 16-bit stereo WAV audio. The model is available on Hugging Face and supports SGLang-Omni, diffusers, and ComfyUI.

Aug 4

Aug 4Tue
  1. Hugging FaceAI score20

    Hugging Face joins Open Secure Alliance on security incident learning guidelines

    AIHugging Face is working with the Open Secure Alliance to develop guidelines for incident learning. The goal is to collectively improve how security incidents are reviewed, disclosed, and controlled. The Alliance, now over 120 members, is sharing proposed SAFE guidelines for turning confidential incident findings into broader ecosystem protection.

Jul 29

Jul 29Wed
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases UEmbed-9B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-9B, a decoder-only multimodal embedding model built on Qwen3.5 9B that outputs both dense and SPLADE-style sparse embeddings from one forward pass. It supports text, image, video, and mixed-modal inputs for retrieval and multimodal search, and the family also includes 2B and 4B variants. The model is available on Hugging Face, with transformers and vLLM inference support.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba NLP releases UEmbed-4B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-4B, a decoder-only multimodal embedding model built on Qwen3.5 4B that outputs both dense and sparse embeddings from one forward pass. It handles text, image, video, and mixed-modal inputs for retrieval and visual-document search, and sparse activations map to vocabulary terms usable with inverted indexes. The model is available on Hugging Face in a family that also includes 2B and 9B variants.

  3. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score43

    Alibaba-NLP releases UEmbed-2B, a multimodal model producing dense and sparse embeddings

    AIAlibaba-NLP's UEmbed-2B, a decoder-only multimodal embedding model built on Qwen3.5 2B, produces both dense and SPLADE-style sparse embeddings from a single forward pass. It supports text, image, video, and mixed-modal inputs for retrieval, and the 4B and 9B variants are also available. The team reports state-of-the-art results on the text and agent tracks of MMEB-v3.

Jul 28

Jul 28Tue
  1. Intern Large ModelsAI score62

    Intern Large Models introduces Visual Pretraining learned from visual documents

    AIIntern Large Models introduces Visual Pretraining, a pretraining paradigm for foundation models that learns directly from visual documents. The post says it outperforms text-only pretraining across backbones and benchmarks, and links the arXiv paper 2607.09657 along with Intern-S2-Preview (35B) and Intern-S2-Preview-397B on Hugging Face, the latter presented as a multimodal foundation model trained with this recipe.

Jul 27

Jul 27Mon
  1. Andrew NgAI score34

    Andrew Ng urges open models for AI defense, rejecting closed-model safety claims

    AIAndrew Ng praised Nvidia's letter and argued that open models and harnesses are needed for defense, citing the OpenAI-Hugging Face hack. He said claims that closed models are safer are regulatory capture. Jensen Huang's background post says closed AI blocked forensics during the Hugging Face incident, while an open-weight frontier model helped contain it, leading to the Open Secure AI Alliance.

  2. KimiAI score86

    Moonshot AI releases Kimi K3 weights and technical report

    AIMoonshot AI is releasing the model weights and technical report for Kimi K3, a 2.8T-parameter MoE model with native visual understanding and a 1M-token context window. The post says the new architecture delivers 2.5x the intelligence per unit of compute, and the company is also opening high-performance attention kernels, an MoE communication library, and infrastructure for running agent environments at scale.

    Why it matters: The source names the model size, context window, and released weights, which helps readers compare its scale and openness with other frontier releases.

Jul 26

Jul 26Sun

Jul 23

Jul 23Thu

Jul 21

Jul 21Tue
  1. Soumith ChintalaAI score45

    Soumith Chintala says Poolside's Laguna S 2.1 suits agentic work on DGX Spark

    AISoumith Chintala praised Poolside's Laguna S 2.1 as looking strong for agentic use and said it fits on a single NVIDIA DGX Spark. The quoted Poolside release describes it as a 118B total-parameter Mixture-of-Experts model with 8B active per token, up to 1M-token context, and thinking and no-thinking modes, with weights openly available under OpenMDW-1.1.