Skip to contentSkip to stories

Updated

#arXiv

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

Oct 9Fri
  1. Rohan PaulXAI score47

    NYU and Amazon paper: keeping a few skills beats distilling a large bank

    AIA New NYU and Amazon paper finds that distilling only the skills that keep giving a useful training signal matches or beats distilling a skill bank up to 11 times larger. The method, SGUID, keeps skills that help early and late in training, and with 6 such skills, 3 of 4 models matched or beat the full bank of 30 to 71 skills on math contest tests. A second round with 3 new skills raised Qwen3-8B from 64.3% to 66.3%.

    Image from @rohanpaul_ai's post
  2. X.PINXAI score46

    Seed preprint finds DeepSeek V4 long-context retrieval varies by position

    AIA Seed team preprint reports "phase sensitivity" in DeepSeek V4 and V4.1-Flash, where identical information becomes harder to retrieve depending on its position within compressed KV-cache blocks. The compression reduces memory and attention costs, but long-context retrieval accuracy varied by up to 40 percentage points across positions. The authors note that average benchmark scores can hide these recurring weak spots, though the findings concern retrieval specifically rather than all model behavior.

    Image from @thexpin's post

Oct 7

Oct 7Wed
  1. Google ResearchOfficialAI score23

    Google Research invites COLM visitors to ContinuousBench walkthrough on DP synthetic data

    AIGoogle Research is hosting a walkthrough at its COLM booth #107 today at 5:00 PM of ContinuousBench, a standardized benchmark for measuring knowledge transfer in differentially private synthetic data. The session, led by Alex Bie, asks whether DP synthetic data preserve actual information or only style. A paper is linked on arXiv.

    Image from @GoogleResearch's post

Oct 6

Oct 6Tue

Oct 1

Oct 1Thu