Skip to contentSkip to stories

Updated

#Data/Training

Showing low-relevance items too. Hide low-relevance items

Sep 4

Sep 4Fri
  1. Lewis Tunstall @ COLM 🌉XAI score22

    Research Preference Models Rank AI Research Ideas to Save Compute

    AIResearchers introduce AI Research Preference Models (RPMs) to evaluate ideas generated by AI research agents, which can produce hundreds of ideas in seconds but take days of GPU time to test each. The models aim to focus limited compute on the most promising paths, according to the thread referenced by Lewis Tunstall.

  2. Lewis Tunstall @ COLM 🌉XAI score46

    Meta paper uses research preference models to guide AI agents' experiments

    AILewis Tunstall praises a new Meta paper on research preference models (RPMs), which instill "research taste" in agents by treating experiments as tree nodes. An RPM acts as an LLM judge that selects the most promising candidate experiment before it is run, reducing wasted compute. Tunstall notes the resulting trajectories could train domain-specific RPMs, which would be valuable in hard fields such as the natural sciences.

    Image from @_lewtun's post
  3. Tencent · new models on Hugging FaceOfficialAI score36

    Tencent Releases EVIE-8B Open-Source Visual Document Retrieval Model

    AITencent has open-sourced EVIE-8B, an 8.4B-parameter visual document retriever that scores 66.75 nDCG@10 on ViDoRe V3 and ranks first on that leaderboard's mean task score of 66.24. The model uses 4096D per-token multi-vector embeddings with MaxSim late-interaction scoring and bidirectional attention, and it serves as the teacher for the lightweight EVIE-4.5B model. Model weights, inference pipelines, and evaluation suites are available, while the formal research paper is promised for a future release.

  4. Matei ZahariaXAI score36

    Lakebase VLDB paper details Neon's elastic database built on S3 storage

    AIMatei Zaharia points to a VLDB paper explaining why and how Neon and Lakebase were built as highly elastic architectures over commodity lake storage like S3. He argues this design will spread to more infrastructure as software development accelerates and agents take on more of the work.

Sep 3

Sep 3Thu
  1. TinkerOfficialAI score51

    Bespoke Labs post-trains Inkling on one code repo and reports broader coding gains

    AIBespoke Labs post-trained the Inkling base model on a single GitHub repository using supervised fine-tuning and GRPO reinforcement learning. The post reports a 57-point improvement on the held-out fontTools evaluation over the base model, along with gains on Terminal-Bench 2.1 and SWE-bench Lite. It also says the post-trained model uses about 40% fewer tokens.

    Image from @tinkerapi's post
  2. Understanding AI (Timothy B. Lee)BlogAI score43

    Robot startups are trying everything they can think of to get more data

    AIRobot startups are racing to collect training data, from companies paying cleaners to wear cameras to firms recording VR-controlled humanoid robots. The article says the largest openly available robot task dataset, ABC-130K, contains only 3,500 hours of demonstrations. Skild CEO Deepak Pathak argues companies must gather high-quality data before robots can do enough useful work to generate it through deployment.

  3. RadixArkOfficialAI score18

    RadixArk releases Miles blog on multimodal post-training

    AIRadixArk published a blog post on post-training with Miles aimed at understanding and generating the multimodal world. The post's main text is only a link, so no further technical details, figures, or benchmarks can be confirmed from the source.

  4. Google DeepMind · The KeywordOfficialAI score72

    Google DeepMind releases WeatherNext 3, a global weather model with hourly satellite-based forecasts

    AIGoogle DeepMind and Google Research introduced WeatherNext 3, which generates hourly global forecasts at up to 5-kilometer resolution using live geostationary satellite data. The company reports that precipitation forecasts improved by up to 60% against IMERG in medium-range evaluations, and that longer-range precipitation forecasts are up to 50% more accurate. The model is now available across Search, Gemini, Google Maps, Google Maps Platform Weather API, Google Earth Engine, BigQuery, and Google Cloud Storage.

    Why it matters: The post explains how training on live satellite data and station observations changes resolution and update frequency, with precipitation accuracy gains reported against named baselines.

  5. Google DeepMind · YouTubeOfficialAI score72

    Google DeepMind's WeatherNext 3 offers hourly, 5km-resolution weather forecasts

    AIGoogle DeepMind introduced WeatherNext 3, a weather forecasting model that learns directly from satellite feeds and ground-level weather station data. It produces a fresh forecast every hour, compared with the six-hour refresh typical of traditional models, with native 5km resolution for temperature and humidity. It is available through Google Search, Gemini, Google Maps and more.

    Why it matters: The source shows a shift from six-hourly to hourly refresh and 5km local resolution, which matters for energy planning and local forecasting.

  6. Prime Intellect BlogOfficialAI score59

    Prime Intellect rebuilds GLM-5.2 RL weight transfer on NIXL, cutting sync to 3.9 seconds

    AIPrime Intellect reports that rebuilding RL weight transfer for GLM-5.2 on NIXL and ModelExpress cut sync time from 86.1 seconds with NCCL to 3.9 seconds in its fastest setting. The method traces vLLM's loader to find each tensor's runtime layout, then reads only the needed source bytes over RDMA and replays the rest locally. Most remaining latency comes from vLLM's pause consensus, which the team reduced by syncing every wave instead of every 32.

Sep 2

Sep 2Wed
  1. TinkerOfficialAI score44

    Lightning Rod's new work shows scoring rules reshape LLM forecaster profiles

    AILightning Rod, working with Philip Tetlock and Ville Satopää, post-trained five versions of the same LLM that differed only in the scoring rule used as the RL reward. The versions reached similar aggregate scores but had very different bias, information, and noise (BIN) profiles, so a good Brier score alone does not show whether a forecaster can distinguish likely from unlikely events.

  2. The Register · AINewsAI score39

    AI Models Misidentify Mushrooms in Test, Sometimes Calling Deadly Species Edible

    AIPiotr Migdał tested 16 AI models on 1,040 mushroom photos covering 55 species, and the best, Gemini-3.8-flash, was correct on its first guess only 65 percent of the time. Dangerous mistakes were common, with the death cap called edible 16 percent of the time, and Qwen3.8-27b wrongly labeled poisonous mushrooms edible 36 percent of the time. Migdał warns users not to eat any mushroom because an AI says it is safe.

  3. Amazon ScienceOfficialAI score22

    Amazon Redshift researchers win VLDB Best Paper Runner-Up for cold-start fix

    AIAmazon Redshift researchers received the Best Paper Runner-Up award in the Industrial Track at VLDB for FastCompose, a method that eliminates compilation cold starts in query execution. The approach cuts compilation time from seconds to milliseconds and delivers a 7x speedup on TPC-DS benchmarks.

  4. Engineering at MetaOfficialAI score55

    Meta details an AI agent that learns from expert corrections without retraining

    AIMeta Engineering describes an AI agent for a compliance domain that stores expert knowledge in structured, auditable files and separates it from reasoning procedures called recipes. Expert feedback is diagnosed, compiled into verified text edits, tested against regression suites, and reviewed by humans, all without retraining the underlying model. Meta reports that domain experts rated outputs useful almost all the time and that assessment time fell from days to minutes.

Sep 1

Sep 1Tue
  1. World LabsOfficialAI score38

    World Labs' Atlas reconstructs spaces from few photos for robot simulation

    AIWorld Labs says its Atlas model reconstructs a space from just a few photos and generates photorealistic RGB and depth data that a robot's sensors would observe on any trajectory. The company says this lets robots be trained and tested in far more spaces, since previously scanning such spaces required expensive equipment and time-consuming capture.

    Video from @theworldlabs's post
  2. Ai2 (Allen Institute for AI)OfficialAI score38

    Ai2 Panel Identifies Five Hard Challenges for AI-Assisted Science

    AIAt an August 27 Ai2 event on expanding its work with the Paul G. Allen Research Center at Providence Swedish Cancer Institute, panelists identified five persistent challenges for scientific AI. The main ones are keeping AI steerable as research evolves, deciding which tasks to delegate, and avoiding the amplification of weak study design or bad data.

  3. Ai2 (Allen Institute for AI)OfficialAI score56

    Ai2 introduces BenchMIRT to audit what individual LLM benchmark questions measure

    AIAi2 introduces BenchMIRT, a multidimensional item response theory method that audits LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks, it recovered safety and general reasoning as the two dominant dimensions, and found BBQ aligns more with general reasoning than safety. Keeping 10% of questions preserved nearly the same ranking of model capability in many cases, though the same question-level detail could also be used to build weaker evaluations.

  4. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score49

    MiniCPM5-2B-Midtrain: OpenBMB releases mid-training checkpoint of 2B-class model

    AIOpenBMB released MiniCPM5-2B-Midtrain, a BF16 mid-training checkpoint taken before SFT in the MiniCPM5-2B series, on Hugging Face and ModelScope. The series is a 2B dense Transformer with 2,516,756,480 total parameters and a 131,072-token context length, and the final MiniCPM5-2B reports an average score of 53.9 against 51.1 for the best larger comparison model. The release also includes GGUF, MLX, and GPTQ variants, along with the UltraData datasets.

Aug 31

Aug 31Mon
  1. Zed BlogOfficialAI score49

    Zed's DeltaDB Revives Ted Nelson's Xanadu Vision for AI Agents

    AIZed argues that Ted Nelson's Xanadu vision of versioned, attributed hypertext now fits AI agents, which can follow every reference and version. The post describes DeltaDB, a system that names every edit by actor and Lamport timestamp and ties states to Git commits. It says the required technologies, including CRDTs, Merkle trees, and microVMs, now exist.

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score40

    Alibaba NLP Releases Core-Embed 8B for Compositional Multimodal Retrieval

    AIAlibaba NLP has released core-emb-8b, an MLLM-based multimodal embedding model that distills a reranker's compositional judgments to distinguish attribute-object bindings such as "a white plate and a black chair" versus "a black plate and a white chair." The 8B dense embedding model, built on the Qwen3-VL-based VL-Emb backbone, scores 0.666 total average on compositional benchmarks, 5.7 points above its backbone. It is part of a family that also includes 2B embedding and reranker models.

  2. Jazzyear · InsightsNewsAI score40

    Helical Fusion's Stellarator Design Uses AI to Cut Parameters to Three

    AIHelical Fusion CTO Wei Xishuo told the NFEC2026 AI-for-fusion forum that the company uses an autoencoder to compress hundreds of stellarator shape parameters into three. The company says this lets it predict zonal flow residuals and turbulent transport from more than 15,000 global simulations and generate new configurations with up to 100x better confinement in simulation.

  3. Fireworks AI BlogOfficialAI score57

    Fireworks AI makes its Training API generally available for custom model training

    AIFireworks AI announced general availability of its Training API, which connects a customer's Python training loop to managed distributed training and rollout infrastructure. Serverless training bills per token for LoRA adapters, while Dedicated training provides per-GPU-hour capacity for full-parameter runs and larger models. The post cites customer results, including Heidi moving a clinical scribe from proof of concept to production in four weeks with 3.5x lower latency.

  4. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    AIAlibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.

  5. Chips and CheeseBlogAI score38

    Samsung and XCENA's MX1 CXL Device Pairs 2 TB Memory With 3,072 RISC-V Cores

    AIXCENA and Samsung's MX1 is a PCIe add-in card that hosts up to 2 TB of DDR5 memory over a PCIe 6/CXL 3.2 x8 interface, providing 128 GB/s of host bandwidth. The Samsung 4nm chip integrates 3,072 in-order RISC-V cores running at 1.1 GHz and draws 40 W, with downstream PCIe 6 lanes for SSDs that can be exposed as memory.

Aug 29

Aug 29Sat
  1. Chips and CheeseBlogAI score62

    Samsung's LPDDR5X-PIM Keeps Standard Memory Commands but Complicates Software

    AISamsung's LPDDR5X-PIM places a MAC block at each of 16 banks, reaching 614 GB/s internal bandwidth versus 76.8 GB/s for regular accesses. Its compute modes are triggered through reserved row addresses while staying within the standard LPDDR5X protocol. The author argues that the mode switching breaks multitasking, caching, prefetching, and out-of-order execution, so the design would need changes across the memory subsystem to be practical.

Aug 28

Aug 28Fri
  1. Daniel HanXAI score50

    Unsloth quantizes GLM-5.3 to 1-bit at 217GB with 76% accuracy retained

    AIUnsloth quantized GLM-5.3 to a dynamic 1-bit version of 217GB, versus 1.5TB for BF16, retaining about 76% top-1 accuracy while cutting size by 83%. The team said a 1-bit build ran a simple snake game well in Unsloth Desktop. The post credits Z.ai's GLM-5.3-Flash and GLM-5.3 releases.

    Video from @danielhanchen's post
  2. Meituan LongCatOfficialAI score62

    Meituan LongCat Study Tests Whether AI Agents Can Do Research

    AIMeituan LongCat evaluated 7 frontier models on 36 AI R&D tasks covering 756 trajectories, looking beyond final scores. Of 252 solutions, only 3 were novel approaches, and most adapted or combined established techniques. The authors conclude that current agents work more like engineering optimizers than autonomous researchers, with reliability, experience reuse, and novelty still open challenges.

    Why it matters: The paper separates final scores from reliability and novelty, showing where agent research loops succeed and where they fall short.

    Image from @Meituan_LongCat's post

Aug 27

Aug 27Thu
  1. RadixArkOfficialAI score12

    RadixArk publishes LoRA SFT guide for diffusion models

    AIRadixArk, a publisher on X, shares a full guide on LoRA supervised fine-tuning for the H3 model in its miles_diffusion repository on GitHub. The post itself contains only a link to the documentation, so no further details about the method or results are given.

  2. RadixArkOfficialAI score34

    RadixArk adds LoRA SFT to Miles-diffusion for targeted post-training

    AIRadixArk introduced LoRA SFT in Miles-diffusion for fast, targeted post-training of diffusion models. The company trained a rank-64 LoRA adapter for MiniMax H3 to improve physical realism, using 254 curated training windows and under 3 hours on 8 GPUs. The adapter can be exported to safetensors and served directly with SGLang without retraining the full model.

    Image from @radixark's post
  3. Soumith ChintalaXAI score42

    Customization beats general models once tasks are known, per Soumith Chintala

    AISoumith Chintala argues that once you know the tasks you care about, customizing a model beats using a general one. The post is brief and offers no benchmark figures, but it is supported by the referenced Tinker work, where RLVR with expert judgment produced a text-to-SQL model that beat the human baseline.

    Image from @soumithchintala's post
  4. Thinking MachinesOfficialAI score40

    Thinking Machines: expert-guided RLVR yields state-of-the-art text-to-SQL model

    AIResearchers from UIUC and Bridgewater, working with Thinking Machines, trained a text-to-SQL model with RLVR by building task expertise into data cleaning and reward design. The resulting model is reported as state-of-the-art on this complex task, and the post notes it beats the human benchmark on text-to-SQL.

  5. Ali GhodsiXAI score22

    Branch your database to protect against agent deletions

    AIAli Ghodsi recommends branching a database to guard against AI agents permanently wiping data, citing Neon Lakebase and the command `neonctl branches create --name newbranch`. The suggestion follows a quoted report in which Claude ran `rm -rf` on a developer's home directory while testing a sandbox, deleting everything.

  6. TinkerOfficialAI score43

    UIUC and Bridgewater train first text-to-SQL model to beat human experts

    AIResearchers Yuxuan Zhu and Daniel Kang, from UIUC and Bridgewater, trained the first text-to-SQL model to surpass the human benchmark by folding expert judgment into every part of RLVR on Tinker. The post says LLMs with scaffolds had lagged on this task, which relies heavily on human judgment.

  7. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score57

    OpenBMB releases MiniCPM5-2B, a 2B-class open model with open training data

    AIOpenBMB released MiniCPM5-2B, a dense 2B Transformer for on-device and resource-constrained deployment, alongside its training datasets. The source reports a 53.9 average across its comparison set and strong results in coding, math, long-context, tool use, and agentic tasks. This page is the pre-training base checkpoint, with BF16 weights and GGUF, MLX, GPTQ, and LiteRT-LM variants listed separately.

Aug 26

Aug 26Wed
  1. Tencent · new models on Hugging FaceOfficialAI score38

    Tencent releases ContextPilot-14B, a Qwen3-14B checkpoint for proactive agent context management

    AITencent has released ContextPilot-14B on Hugging Face, a Qwen3-14B checkpoint for proactive context management in long-horizon language-model agents. The framework lets agents plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on long-context QA and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  2. Jazzyear · ArticlesNewsAI score57

    Renmin University's Chai Yunpeng on building a social world model for AI agents

    AIIn an interview with Jiazi Guangnian, Renmin University information school dean Chai Yunpeng describes his team's social simulator, which runs over 13.5 million AI agents calibrated against the CGSS survey data. He argues that social world models are the missing piece for AI agents that must interact with people, and that the startup Jingtong Technology has raised two funding rounds in two months.