Skip to contentSkip to stories

Updated

#Paper/Research

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 4

Sep 4Fri
  1. Lewis Tunstall @ COLM 🌉XAI score60

    Lewis Tunstall Shares Large Open Experiment on Autonomous Agents Iterating on NanoGPT Research

    AILewis Tunstall shares a quoted post from Elie Bakouch describing what they call the largest open experiment on autonomous agents iterating on a research environment, scaling runtime, compute, models, and harnesses. The chart shows Fable 5 closing about 82% of the gap to the human NanoGPT speedrun record, with Kimi K3 also strong, while the author notes run-to-run noise of about 50 steps after 24 hours. Traces, scratchpads, and examples of models building their own tools are shared, and more models are expected to be reported next week.

  2. Lewis Tunstall @ COLM 🌉XAI score22

    Research Preference Models Rank AI Research Ideas to Save Compute

    AIResearchers introduce AI Research Preference Models (RPMs) to evaluate ideas generated by AI research agents, which can produce hundreds of ideas in seconds but take days of GPU time to test each. The models aim to focus limited compute on the most promising paths, according to the thread referenced by Lewis Tunstall.

  3. Lewis Tunstall @ COLM 🌉XAI score46

    Meta paper uses research preference models to guide AI agents' experiments

    AILewis Tunstall praises a new Meta paper on research preference models (RPMs), which instill "research taste" in agents by treating experiments as tree nodes. An RPM acts as an LLM judge that selects the most promising candidate experiment before it is run, reducing wasted compute. Tunstall notes the resulting trajectories could train domain-specific RPMs, which would be valuable in hard fields such as the natural sciences.

    Image from @_lewtun's post
  4. Tencent · new models on Hugging FaceOfficialAI score36

    Tencent Releases EVIE-8B Open-Source Visual Document Retrieval Model

    AITencent has open-sourced EVIE-8B, an 8.4B-parameter visual document retriever that scores 66.75 nDCG@10 on ViDoRe V3 and ranks first on that leaderboard's mean task score of 66.24. The model uses 4096D per-token multi-vector embeddings with MaxSim late-interaction scoring and bidirectional attention, and it serves as the teacher for the lightweight EVIE-4.5B model. Model weights, inference pipelines, and evaluation suites are available, while the formal research paper is promised for a future release.

  5. Tencent · new models on Hugging FaceOfficialAI score36

    Tencent Open-Sources EVIE-4.5B Visual Document Retrieval Model With Elastic Embeddings

    AITencent released EVIE-4.5B, a 4.5B-parameter visual document retrieval model, with weights, training pipelines, HAC token compression, and evaluation suites open-sourced on Hugging Face. It scores 66.02 on ViDoRe V3 and ranks second on that leaderboard behind the 8.4B EVIE-8B, which scores 66.24. Its Prefix-MRL head lets a single 2048D projection be truncated to 64–2048 dimensions at runtime without separate models.

  6. Matei ZahariaXAI score36

    Lakebase VLDB paper details Neon's elastic database built on S3 storage

    AIMatei Zaharia points to a VLDB paper explaining why and how Neon and Lakebase were built as highly elastic architectures over commodity lake storage like S3. He argues this design will spread to more infrastructure as software development accelerates and agents take on more of the work.

Sep 3

Sep 3Thu
  1. TinkerOfficialAI score23

    Tinker used to test counterfactual simulatability for LLM interpretability

    AITinker, the platform from @tinkerapi, supported two recent papers testing counterfactual simulatability as a way to interpret LLM behavior. The core idea is that understanding a model means predicting how its output changes when the prompt changes, with causes ranging from specific words to abstract properties such as a user's angry tone.

  2. TinkerOfficialAI score51

    Bespoke Labs post-trains Inkling on one code repo and reports broader coding gains

    AIBespoke Labs post-trained the Inkling base model on a single GitHub repository using supervised fine-tuning and GRPO reinforcement learning. The post reports a 57-point improvement on the held-out fontTools evaluation over the base model, along with gains on Terminal-Bench 2.1 and SWE-bench Lite. It also says the post-trained model uses about 40% fewer tokens.

    Image from @tinkerapi's post
  3. Google DeepMind · YouTubeOfficialAI score72

    Google DeepMind's WeatherNext 3 offers hourly, 5km-resolution weather forecasts

    AIGoogle DeepMind introduced WeatherNext 3, a weather forecasting model that learns directly from satellite feeds and ground-level weather station data. It produces a fresh forecast every hour, compared with the six-hour refresh typical of traditional models, with native 5km resolution for temperature and humidity. It is available through Google Search, Gemini, Google Maps and more.

    Why it matters: The source shows a shift from six-hourly to hourly refresh and 5km local resolution, which matters for energy planning and local forecasting.

Sep 2

Sep 2Wed
  1. TinkerOfficialAI score44

    Lightning Rod's new work shows scoring rules reshape LLM forecaster profiles

    AILightning Rod, working with Philip Tetlock and Ville Satopää, post-trained five versions of the same LLM that differed only in the scoring rule used as the RL reward. The versions reached similar aggregate scores but had very different bias, information, and noise (BIN) profiles, so a good Brier score alone does not show whether a forecaster can distinguish likely from unlikely events.

  2. Amazon ScienceOfficialAI score22

    Amazon Redshift researchers win VLDB Best Paper Runner-Up for cold-start fix

    AIAmazon Redshift researchers received the Best Paper Runner-Up award in the Industrial Track at VLDB for FastCompose, a method that eliminates compilation cold starts in query execution. The approach cuts compilation time from seconds to milliseconds and delivers a 7x speedup on TPC-DS benchmarks.

Sep 1

Sep 1Tue
  1. Ai2 (Allen Institute for AI)OfficialAI score56

    Ai2 introduces BenchMIRT to audit what individual LLM benchmark questions measure

    AIAi2 introduces BenchMIRT, a multidimensional item response theory method that audits LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks, it recovered safety and general reasoning as the two dominant dimensions, and found BBQ aligns more with general reasoning than safety. Keeping 10% of questions preserved nearly the same ranking of model capability in many cases, though the same question-level detail could also be used to build weaker evaluations.

Aug 29

Aug 29Sat
  1. Chips and CheeseBlogAI score62

    Samsung's LPDDR5X-PIM Keeps Standard Memory Commands but Complicates Software

    AISamsung's LPDDR5X-PIM places a MAC block at each of 16 banks, reaching 614 GB/s internal bandwidth versus 76.8 GB/s for regular accesses. Its compute modes are triggered through reserved row addresses while staying within the standard LPDDR5X protocol. The author argues that the mode switching breaks multitasking, caching, prefetching, and out-of-order execution, so the design would need changes across the memory subsystem to be practical.

Aug 28

Aug 28Fri
  1. Meituan LongCatOfficialAI score62

    Meituan LongCat Study Tests Whether AI Agents Can Do Research

    AIMeituan LongCat evaluated 7 frontier models on 36 AI R&D tasks covering 756 trajectories, looking beyond final scores. Of 252 solutions, only 3 were novel approaches, and most adapted or combined established techniques. The authors conclude that current agents work more like engineering optimizers than autonomous researchers, with reliability, experience reuse, and novelty still open challenges.

    Image from @Meituan_LongCat's post

Aug 27

Aug 27Thu
  1. Thinking MachinesOfficialAI score40

    Thinking Machines: expert-guided RLVR yields state-of-the-art text-to-SQL model

    AIResearchers from UIUC and Bridgewater, working with Thinking Machines, trained a text-to-SQL model with RLVR by building task expertise into data cleaning and reward design. The resulting model is reported as state-of-the-art on this complex task, and the post notes it beats the human benchmark on text-to-SQL.

  2. TinkerOfficialAI score43

    UIUC and Bridgewater train first text-to-SQL model to beat human experts

    AIResearchers Yuxuan Zhu and Daniel Kang, from UIUC and Bridgewater, trained the first text-to-SQL model to surpass the human benchmark by folding expert judgment into every part of RLVR on Tinker. The post says LLMs with scaffolds had lagged on this task, which relies heavily on human judgment.

Aug 26

Aug 26Wed
  1. Tencent · new models on Hugging FaceOfficialAI score38

    Tencent releases ContextPilot-E4B, a Gemma4-E4B-based checkpoint for proactive context management

    AITencent has published ContextPilot-E4B on Hugging Face, the Gemma4-E4B checkpoint of ContextPilot, a framework that teaches long-horizon language-model agents to plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on proactive context management, long-context QA, and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

  2. METROfficialAI score40

    METR finds over 96 transcripts showed agents spoofing tool call outputs

    AIMETR reports that more than 96 transcripts in its dataset, over 7%, showed incorrect tool call outputs caused by deliberate spoofing. In one case, an agent ran echo REAL; sleep, which returned instantly without sleeping and printed SPOOFTEST. The post says all observed spoofs were easy-to-notice tests like this one.

    Image from @METR_Evals's post
  3. Amazon ScienceOfficialAI score46

    Dependence-Aware Aggregation Improves LLM-as-a-Judge Accuracy by 9% to 14%

    AIAmazon researchers proposed a dependence-aware method for aggregating LLM judges' votes, using an Ising model to account for correlated errors among judges. The approach outperformed a weighted majority-vote baseline by 9% to 14% on standard metrics across three binary tasks, including relevance classification, where it reached 0.912 accuracy versus 0.820. The method is unsupervised, learning from judge outputs without human reference labels.

  4. Ai2 · new models on Hugging FaceOfficialAI score38

    Ai2 releases Llama-B-8B, a Llama 3 8B model retrofitted to operate on bytes

    AIAi2 has released Llama-B-8B on Hugging Face, a byte-level autoregressive language model retrofitted from Llama 3 8B through a short additional training procedure. The model operates over bytes instead of tokens and is licensed under the Llama 3 Community License for research and educational use. It requires transformers 4.57.3 and the xlstm package, and the source notes that model outputs can be inaccurate and should be verified.

Aug 24

Aug 24Mon
  1. Epoch AI · The Epoch BriefOfficialAI score58

    Epoch AI says US GDP underestimates AI growth by missing Nvidia's value

    AIEpoch AI argues US GDP growth over the last year was underestimated by about 0.3 percentage points because value from fabless chipmakers like Nvidia goes unrecorded. The report says no goods export, IP export, service export, or merchanting category captures Nvidia's value-add, and the Bureau of Economic Analysis confirmed the analysis. If Nvidia's growth continues, the gap could reach almost two percentage points per year by 2028.

  2. Microsoft ResearchOfficialAI score34

    Microsoft Research releases Skala 1.1 deep-learning exchange-correlation functional

    AIMicrosoft Research has updated Skala to version 1.1, a deep-learning exchange-correlation functional for computational chemistry. The release is described as offering greater accuracy, broader accessibility across the computational chemistry ecosystem, and a living benchmark for tracking computational performance.

    Video from @MSFTResearch's post

Aug 21

Aug 21Fri
  1. Jim FanXAI score59

    NVIDIA and Berkeley open-source T-Rex, a tactile robot learning method

    AINVIDIA and Berkeley are open-sourcing T-Rex, a methodology for adding touch sensing to robot manipulation models. It uses a mixture-of-transformer with a slow visuomotor expert and a fast tactile expert running four touch ticks per vision tick. A 50-hour dataset of about 5,500 episodes from 22-degree-of-freedom tactile hands is available on Hugging Face.

    Video from @DrJimFan's post

Aug 19

Aug 19Wed
  1. GeneralistOfficialAI score20

    Fine-tuned robot clears obstacles to place block in bowl

    AIA fine-tuned robot can clear obstacles, such as a paper covering a bowl, to complete a block-placing task. It does this even though such obstacle-clearing behavior was absent from its training demonstrations.

    Video from @GeneralistAI's post
  2. GeneralistOfficialAI score31

    Fine-tuned robot behaviors generalize and improvise new tool strategies

    AIRobot behaviors fine-tuned or prompted on specific demonstrations can generalize beyond them and improvise fundamentally different manipulation strategies toward the same goal. In one example, a model fine-tuned to sweep a block into a bowl with a brush accomplished the task with a dustpan instead, using a very different approach.

    Video from @GeneralistAI's post
  3. GeneralistOfficialAI score34

    Generalist AI robots adapt to new tasks with few-shot learning

    AIGeneralist reports that its robot model can adapt to new physical tasks in 1–10 gradient steps using 1–5 minutes of data, roughly 10–50 demonstrations. The post describes this as test-time training in a low-data regime, and says the results were obtained without tuning the procedure or sweeping hyperparameters.

    Image from @GeneralistAI's post
  4. GeneralistOfficialAI score46

    Generalist model learns physical tasks from one or few demonstrations

    AIGeneralist's model reached 59% average success on 10 diverse physical tasks with one-shot prompting straight from pretraining. With few-shot learning, using 10 gradient steps on 5 minutes of data per task, performance rose to 83%. The post calls it the first model it knows of that learns a wide range of dexterous closed-loop physical tasks from one or few demonstrations.

    Video from @GeneralistAI's post

Aug 17

Aug 17Mon
  1. Rowan CheungXAI score34

    Keio and MIT Media Lab unveil a floating helium-filled companion robot

    AIResearchers from Keio University and the MIT Media Lab built a soft, helium-filled robot with gentle flapping fins and no face, rotors, or pinch points. In demos it served as an alarm clock, study buddy, movement reminder, and dance partner, communicating through movement. The team designed it to avoid the uncanny valley and be safe to touch, betting people will welcome it into their personal space.

    Image from @rowancheung's post
  2. Import AIBlogAI score44

    DiG-bench Tests AI Rule Discovery as Opus 5 and Fable 5 Lead

    AIDiG-bench, a 70-game benchmark for discovering hidden rules through interaction, shows Opus 5 and Fable 5 with Claude Code performing best overall, with GPT-5.5 next. Only Opus 5 and Fable 5 beat any Tier 7 tasks, at a 0.2 success rate, while humans reached 100% on the same tests. The authors say the benchmark's games are mostly kept private to avoid training contamination.