Skip to contentSkip to stories

Updated

#Data/Training

Showing low-relevance items too. Hide low-relevance items

Aug 26

Aug 26Wed
  1. Amazon ScienceOfficialAI score46

    Dependence-Aware Aggregation Improves LLM-as-a-Judge Accuracy by 9% to 14%

    AIAmazon researchers proposed a dependence-aware method for aggregating LLM judges' votes, using an Ising model to account for correlated errors among judges. The approach outperformed a weighted majority-vote baseline by 9% to 14% on standard metrics across three binary tasks, including relevance classification, where it reached 0.912 accuracy versus 0.820. The method is unsupervised, learning from judge outputs without human reference labels.

  2. Ai2 · new models on Hugging FaceOfficialAI score38

    Ai2 releases Llama-B-8B, a Llama 3 8B model retrofitted to operate on bytes

    AIAi2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.

  3. Ai2 · new models on Hugging FaceOfficialAI score37

    Ai2 releases Llama-B 8B Stage 1 checkpoint, a byte-level Llama 3 8B variant

    AIAi2 has released allenai/Llama-B-8B-Stage1, a Llama 3 8B model retrofitted to operate over bytes instead of tokens through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 or later and the xlstm package, and is loaded with trust_remote_code.

  4. Ai2 · new models on Hugging FaceOfficialAI score38

    Ai2 releases Bwen-8B, a byte-level model retrofitted from Qwen3 8B Base

    AIAi2 has released Bwen-8B, a byte-level autoregressive language model retrofitted from Qwen3 8B Base through a short additional training procedure called byteification, which lets it operate over bytes instead of tokens. The model is licensed under Apache 2.0 for research and educational use, and requires transformers 4.57.3 or later and the xlstm package.

  5. Ai2 · new models on Hugging FaceOfficialAI score38

    Ai2 releases Bwen-8B-Stage1, a byte-level Qwen3-8B retrofit under Apache 2.0

    AIAi2 has released Bwen-8B-Stage1 on Hugging Face, a byte-level autoregressive model retrofitted from Qwen3-8B-Base through a short additional training procedure. This Stage 1 checkpoint contains only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.

  6. Ai2 · new models on Hugging FaceOfficialAI score39

    Ai2 releases Bolmo-1B-Stage1, a byte-level version of OLMo 2 1B

    AIAi2 has released Bolmo-1B-Stage1, a 1B-parameter byte-level language model retrofitted from OLMo 2 1B to process bytes rather than tokens. This checkpoint includes Stage 1 training only, with inner model parameters unchanged, and is available on Hugging Face under an Apache 2.0 license for research and educational use.

  7. Ai2 · new models on Hugging FaceOfficialAI score40

    Ai2 Releases Bolmo-7B-Stage1, a Byte-Level Model Retrofitted from Olmo 3 7B

    AIAi2 has released Bolmo-7B-Stage1, a 7B byte-level language model retrofitted from Olmo 3 7B through a short additional training procedure. This checkpoint includes only Stage 1 training, with inner model parameters unchanged, and is licensed under Apache 2.0 for research and educational use.

Aug 25

Aug 25Tue
  1. Google Developers BlogOfficialAI score35

    Google Brings Qwen3-Embedding-8B to Cloud TPU via vLLM with Long-Context Support

    AIGoogle Cloud has added native TPU support to vLLM and engineered optimizations to serve the Qwen3-Embedding-8B model on Cloud TPU, targeting 4K+ token text and 15K+ token multimodal inputs. The work addresses tensor alignment, lazy-loading, compilation pre-warming, and long-context pooling, with a cosine similarity pass threshold of at least 0.999 for text and 0.995 for multimodal inputs against XPU reference vectors.

  2. Fireworks AI BlogOfficialAI score52

    Harvey Tenet, a legal model post-trained from Kimi K3 with Fireworks

    AIHarvey and Fireworks post-trained Tenet from the Kimi K3 base using asynchronous reinforcement learning on the Fireworks Training API for long-horizon legal work. On the Legal Agent Benchmark, Tenet reached 19.7% all-pass versus 10.8% for base Kimi K3, and its cost per task was $5.92 versus $5.62.

  3. Daniel HanXAI score34

    Fine-tune Qwen3.8-27B free on Kaggle with Unsloth QLoRA

    AIDaniel Han says users can fine-tune Qwen3.8-27B for free on Kaggle with a Google account, which provides 30 hours of GPU time on 2× Tesla T4s. Using QLoRA and Unsloth's kernels, the 27B model fits within 24 GB VRAM with no accuracy loss, according to the post. The background post from Unsloth adds that its notebook trains Qwen3.8-27B 1.5x faster with 50% less VRAM.

Aug 24

Aug 24Mon
  1. Google · new models on Hugging FaceOfficialAI score40

    Google releases TimesFM 3.0 time-series forecasting model weights on Hugging Face

    AIGoogle Research has published the official PyTorch weights and configurations for TimesFM 3.0, a pretrained time-series foundation model for forecasting. The model uses a Stacked Mixing Transformer with 20 layers, a model dimension of 1280, and 16 heads, and it is released under the TimesFM Non-Commercial License v1.0.

  2. Epoch AI · The Epoch BriefOfficialAI score58

    Epoch AI says US GDP underestimates AI growth by missing Nvidia's value

    AIEpoch AI argues US GDP growth over the last year was underestimated by about 0.3 percentage points because value from fabless chipmakers like Nvidia goes unrecorded. The report says no goods export, IP export, service export, or merchanting category captures Nvidia's value-add, and the Bureau of Economic Analysis confirmed the analysis. If Nvidia's growth continues, the gap could reach almost two percentage points per year by 2028.

  3. Engineering at MetaOfficialAI score72

    Meta details MetaRoCE, an RDMA transport designed for AI-scale Ethernet

    AIMeta designed MetaRoCE, a clean-sheet RDMA transport for AI workloads on commodity Ethernet, and is releasing its specification, reference software and compliance test suite through the Open Compute Project. On a 64-node AMD GPU cluster running RCCL collectives, the post reports MetaRoCE delivering higher throughput and lower flow completion times than RoCEv2, with about 86% throughput maintained at 1% packet loss.

    Why it matters: The post explains how per-path endpoint intelligence replaces lossless fabric assumptions, with measured throughput and loss results against RoCEv2 on a 64-node AMD cluster.

  4. Microsoft ResearchOfficialAI score34

    Microsoft Research releases Skala 1.1 deep-learning exchange-correlation functional

    AIMicrosoft Research has updated Skala to version 1.1, a deep-learning exchange-correlation functional for computational chemistry. The release is described as offering greater accuracy, broader accessibility across the computational chemistry ecosystem, and a living benchmark for tracking computational performance.

    Video from @MSFTResearch's post
  5. GammaOfficialAI score22

    Gamma's agent usage rose to one in four creators, mostly for edits

    AIGamma says one in four of its creators now use its agent, up from one in twenty a year earlier. Most requests are edits such as fixing charts or shortening text (65.2%), with image generation at 24.5% and creating from scratch at only 4%. Average conversations have grown to 10.2 messages, up 1.6x since March 2025.

    Image from @GammaApp's post
  6. GeneralistOfficialAI score27

    Generalist releases GEN-1.5, a foundation model for physical-world robotics

    AIGeneralist has announced GEN-1.5, its latest foundation model for the physical world. The post provides only a link to the company's blog for further details, so no specifications, benchmarks, or availability information can be confirmed from this source.

  7. GeneralistOfficialAI score38

    Generalist reduces time from physical prompt to robot behavior with GEN-1.5

    AIGeneralist says it has reduced the time needed to go from a physical prompt to robot behavior, making it faster to teach robots new tasks. The company links this speedup to easier scaling of physical work, and points readers to its GEN-1.5 blog post for details.

    Video from @GeneralistAI's post

Aug 23

Aug 23Sun

Aug 21

Aug 21Fri
  1. swyxXAI score39

    Swyx Says Simulating Humans Could Be Last Barrier to Automated AI Research

    AISwyx argues that simulating humans and their feedback is likely the final barrier to recursive self-improvement, where models automate increasingly large parts of ML research. He says Simile, which builds human simulations, is already finding product-market fit with Fortune 100 companies despite its early stage.

  2. Jim FanXAI score59

    NVIDIA and Berkeley open-source T-Rex, a tactile robot learning method

    AINVIDIA and Berkeley are open-sourcing T-Rex, a methodology for adding touch sensing to robot manipulation models. It uses a mixture-of-transformer with a slow visuomotor expert and a fast tactile expert running four touch ticks per vision tick. A 50-hour dataset of about 5,500 episodes from 22-degree-of-freedom tactile hands is available on Hugging Face.

    Video from @DrJimFan's post
  3. Amazon ScienceOfficialAI score50

    SOP-Bench Tests AI Agents on Real Business Procedures Across 12 Industries

    AIAmazon Science released SOP-Bench, an open benchmark that measures how well AI agents execute standard operating procedures written by domain experts. It covers 12 business areas, including healthcare intake and dangerous-goods classification, with more than 2,000 tasks, working tools, and ground-truth answers. The benchmark was presented at the 2026 KDD conference.

Aug 20

Aug 20Thu
  1. Ali GhodsiXAI score20

    Disaggregated storage took off after full bisection bandwidth networks emerged

    AIAli Ghodsi says disaggregating storage from compute became feasible only after research on full bisection bandwidth networks removed datacenter bottlenecks around 2010. Databricks and Snowflake followed soon after, and many others came later. He says putting data on an object store is now the standard approach.

  2. Mistral AIOfficialAI score59

    Mistral Agentic Search adds multi-step retrieval for complex enterprise documents

    AIMistral has released Agentic Search, a multi-step retrieval layer available through its Search Toolkit and Libraries. On FinanceBench, the company reports accuracy rising from 26.7% to 86% over one-shot RAG, and on OfficeQA Pro a gain from 6.3% to 51.9%. The system also reports up to 39.6% lower p90 latency and up to one-third lower token use from fewer repeated searches.

Aug 19

Aug 19Wed
  1. Ali GhodsiXAI score33

    Databricks launches AI Extract for accurate PDF field extraction

    AIDatabricks has launched AI Extract, a capability for extracting fields from PDFs that it says reaches 95% accuracy versus 87% for other tools, at very low cost. The post notes that LLMs' next-token training makes them "autocorrect" content they should preserve, which this approach is designed to avoid. The function can be called directly from SQL and used across the Databricks platform.

    Image from @alighodsi's post
  2. Matei ZahariaXAI score46

    Databricks' custom AI Extract model reaches new frontier in document processing

    AIDatabricks says its in-house AI Extract model, paired with a custom agent harness, achieves a new frontier on complex document processing tasks. The system handles documents over 500 pages and more than 1M tokens, plus nested schemas with 1k+ objects. It decomposes large jobs, runs smaller tasks in parallel, and reconciles them into one structured output.

  3. GeneralistOfficialAI score42

    Generalist's GEN-1.5 marks a new frontier in physical AI generality

    AIGeneralist says its GEN-1.5 model shows a new level of generality when pretrained on physical interaction data at a scale few thought feasible without shortcuts. The company says it does not yet see where performance gains level off.

  4. GeneralistOfficialAI score20

    Fine-tuned robot clears obstacles to place block in bowl

    AIA fine-tuned robot can clear obstacles, such as a paper covering a bowl, to complete a block-placing task. It does this even though such obstacle-clearing behavior was absent from its training demonstrations.

    Video from @GeneralistAI's post
  5. GeneralistOfficialAI score31

    Fine-tuned robot behaviors generalize and improvise new tool strategies

    AIRobot behaviors fine-tuned or prompted on specific demonstrations can generalize beyond them and improvise fundamentally different manipulation strategies toward the same goal. In one example, a model fine-tuned to sweep a block into a bowl with a brush accomplished the task with a dustpan instead, using a very different approach.

    Video from @GeneralistAI's post
  6. GeneralistOfficialAI score34

    Generalist AI robots adapt to new tasks with few-shot learning

    AIGeneralist reports that its robot model can adapt to new physical tasks in 1–10 gradient steps using 1–5 minutes of data, roughly 10–50 demonstrations. The post describes this as test-time training in a low-data regime, and says the results were obtained without tuning the procedure or sweeping hyperparameters.

    Image from @GeneralistAI's post
  7. GeneralistOfficialAI score46

    Generalist model learns physical tasks from one or few demonstrations

    AIGeneralist's model reached 59% average success on 10 diverse physical tasks with one-shot prompting straight from pretraining. With few-shot learning, using 10 gradient steps on 5 minutes of data per task, performance rose to 83%. The post calls it the first model it knows of that learns a wide range of dexterous closed-loop physical tasks from one or few demonstrations.

    Video from @GeneralistAI's post
  8. Google · new models on Hugging FaceOfficialAI score22

    Google releases TIPS g/14 low-res v1 vision-language model on Hugging Face

    AIGoogle has released TIPS g/14 low-res (v1) on Hugging Face, a Text-Image Pre-training with Spatial awareness vision-language model with 1.1B vision parameters and 389M text parameters. The model produces spatially rich image features aligned with text embeddings at 224 resolution, under the Apache 2.0 license. It supports image encoding, text encoding, and zero-shot classification via the transformers library.