Skip to contentSkip to stories

Updated

#Data/Training

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 10

TodayOct 10Sat
  1. 🚨 AI News | TestingCatalogXAI score22

    Hidden Claude Voice tab and voice data sharing option spotted

    AIA hidden Voice tab labeled "Preview" has appeared in Claude's apps, suggesting Anthropic may offer an upgraded voice experience. A new option to share voice data for training has also recently appeared in Claude's mobile apps, but it is unclear whether Anthropic is developing its own voice models.

    Image from @testingcatalog's post
  2. elvisXAI score50

    Sakana AI proposes multi-agent self-supervision for recursive self-improvement without verifiers

    AISakana AI's MASS method lets one base model propose, run, and grade multi-agent workflows, keeping the best through evolutionary search, with no external verifier needed for open-ended tasks. Two self-improvement cycles on Qwen3.6-27B raise performance per output token from 1.2x to 1.6x across four open-ended benchmarks. A student trained on multi-agent traces also beats a single-agent student trained on 1.4x more tokens.

    Image from @omarsar0's post
  3. IThome · AINewsAI score62

    Anthropic's Claude Science helps map the first full-sky ultraviolet survey

    AIAnthropic says Claude Science helped produce the first complete all-sky ultraviolet map by merging data from several space telescope missions. About one third of the map was estimated by the model in regions lacking ultraviolet observations, and the map labels each pixel as measured or predicted with an uncertainty estimate.

  4. Latent SpaceBlogAI score49

    Standard Bots builds AI for reliable industrial robot execution

    AIStandard Bots, which calls itself America's largest AI-native industrial robot manufacturer, runs a shared base model that customers adapt through demonstrations and fine-tuning, with its largest model in the low billions of parameters. CEO Evan Beard and Head of AI Leif Jentoft say the company controls the arm, end effector, control system and AI, and runs inference on-premises on edge GPUs because most factories lack reliable internet.

  5. LeiphoneNewsAI score38

    Li Auto says it will self-develop batteries for all new models from 2026

    AILi Auto says all new models launched from 2026 will use its self-developed batteries, which are already in the new L8, L6 and i8 and will reach the new MEGA, i9 and 2026 i6 in the fourth quarter. Li Auto president Ma Donghui says the move aims to connect data across cell production, pack assembly, vehicle integration and after-sales service. The company says it has about 300 battery R&D staff and nearly 4,000 patent applications and granted patents.

  6. meng shaoXAI score79

    Xiaomi's MiMo-V2.6 report explains scaling RL along batch, environments, and grading

    AIXiaomi's MiMo-V2.6 technical report argues that scaling reinforcement learning, not more pretraining data, is the main lever for frontier capability, along batch size, environment diversity, and grader strength. The article summarizes the report's methods, including groupwise agentic grading, a frozen-router fix for expert load collapse, and reward hacking defenses. It reports RL post-training costs of $2.6 million for MiMo-V2.6-Pro and $0.9 million for MiMo-V2.6-Flash.

    Why it matters: The piece walks through the report's three-way RL scaling method, batch size, environments, and grading, with concrete failure modes and stabilization fixes useful to agentic RL practitioners.

  7. MarkTechPostNewsAI score44

    Microsoft releases Microsoft-Decision-1, a Qwen3.5-9B decision-scoring model

    AIMicrosoft releases Microsoft-Decision-1, a decision-scoring model post-trained from Alibaba's Qwen3.5-9B that returns calibrated probabilities for fixed answer options instead of generated text. The model is available through Microsoft Foundry and OpenRouter at $0.042 per million input tokens, with free output, and has a 32,768-token context window. Microsoft reports 83.5% average accuracy across 36 benchmarks and 85 ms p50 latency, though the results are vendor-run and the latency comparison with competitors is disputed.

  8. Rohan PaulXAI score60

    Xiaomi's MiMo-V2.6 paper details scaling RL for self-improving coding agents

    AIXiaomi's MiMo-V2.6 paper says agents now build tasks, audit tests, grade answers and detect cheating during RL training, with humans setting the budget and rules. A grader agent that rewards cleaner patches over reward-hacking fixes is credited with stopping drift toward longer runs and workarounds such as swallowed exceptions. MiMo-V2.6-Pro's DeepSWE score rose from 58.4 to 72.6 over $2.6M of RL compute and was still climbing when training stopped.

    Image from @rohanpaul_ai's post

Oct 9

Oct 9Fri
  1. Rohan PaulXAI score47

    NYU and Amazon paper: keeping a few skills beats distilling a large bank

    AIA New NYU and Amazon paper finds that distilling only the skills that keep giving a useful training signal matches or beats distilling a skill bank up to 11 times larger. The method, SGUID, keeps skills that help early and late in training, and with 6 such skills, 3 of 4 models matched or beat the full bank of 30 to 71 skills on math contest tests. A second round with 3 new skills raised Qwen3-8B from 64.3% to 66.3%.

    Image from @rohanpaul_ai's post
  2. AMDOfficialAI score22

    AMD and Zyphra train an advanced reasoning model from scratch

    AIAMD's Data Center team says it worked with Zyphra to train an advanced reasoning model from scratch on AMD hardware. A linked post says Zyphra trains larger reasoning models more efficiently while supporting longer context windows.

    Image from @AMD's post
  3. 👩‍💻 Paige BaileyXAI score33

    Encrypted reasoning blocks leak PII and credentials from shared LLM logs

    AIA paper decoded 315,320 reasoning blocks scraped from public repositories and recovered 367 PII artifacts and 182 credentials. The authors say reasoning traces can reveal hazardous information even when the model's visible output refuses a malicious request. They also warn that attackers could hide prompt injections in encrypted blocks to poison public agentic rollouts.

  4. The Robot ReportNewsAI score36

    Boston Dynamics details its redesigned four-fingered humanoid robot hand

    AIBoston Dynamics says its new humanoid hand has four fingers, 13 degrees of freedom, and direct actuation, and it is designed for mass manufacturing. The company dropped the pinky finger and reduced the gripper size, and built the hand for high-fidelity simulation to enable sim-to-real reinforcement learning. Alberto Rodriguez, director of robot behavior for Atlas, says the earlier hands could already lift more than 100 lb., so the new design focuses on reliability and tool use.

  5. Mistral AI · new models on Hugging FaceOfficialAI score32

    Mistral releases LIDstral-Arabic, a language and dialect classifier for Arabic-script text

    AIMistral AI has released LIDstral-Arabic, a fast classifier that identifies Modern Standard Arabic, Arabic dialects, and non-Arabic languages written in Arabic script across 51 classes. On Moroccan Darija, it scores 88.67% F1, versus 72.65% for LahjatBERT ALDi CL and 71.28% for GlotLID v3, across 84,870 evaluation examples. The model runs on CPU and is available from a private Hugging Face repository under Apache 2.0.

  6. Ai2 (Allen Institute for AI)OfficialAI score46

    Ai2 describes GPU time budgets that replaced its priority-based cluster scheduler

    AIAi2's AI Infrastructure team replaced its priority-based scheduler for GPU clusters with GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The team says the change moved debates over how much GPU time each research project deserves from case-by-case operational decisions into a transparent budgeting process. The clusters range from 88 to 1024 GPUs across NVIDIA H100, B200, and B300 hardware, and serve about 150 internal researchers.

  7. Hugging Face BlogOfficialAI score38

    Ai2 replaces priority scheduler with GPU time budgets for cluster allocation

    AIAi2's AI Infrastructure team replaced its priority-based GPU cluster scheduler with a system using GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The team says the change turns decisions about how much GPU time each research project receives into a transparent administrative budgeting process. Its clusters, which range from 88 to 1024 GPUs including H100, B200, and B300 units, serve about 150 researchers facing demand two to three times available capacity.

  8. Baseten BlogOfficialAI score61

    How to choose which layers to run at NVFP4 quantization precision

    AIBaseten explains how to decide which layers of a model can run in 4-bit NVFP4 without losing needed information. The post compares architecture-based heuristics, isolated-layer sensitivity scoring, and SaturationQuant, which accounts for other quantized layers. It also covers calibration with representative data and block-level scales of 16 values.

    Why it matters: The post explains how to choose which layers run at NVFP4 precision using heuristics, sensitivity scoring, and saturation-aware scoring, with clear calibration steps.

  9. DatabricksOfficialAI score25

    Databricks pairs Temporal and Lakebase for durable cloud agents

    AIDatabricks has published a reference implementation pairing Temporal with Lakebase Postgres so cloud agents can survive worker, container, or deployment replacement. The design keeps recorded work and evidence and review state queryable, and lets human decisions arrive days later. Unity Catalog remains the governed policy source through synced tables.

    Image from @databricks's post
  10. Rohan PaulXAI score46

    Microsoft's TeleTune evolves agent skills from raw usage logs

    AIMicrosoft researchers present TeleTune, which lets agents learn software skills from raw usage logs by keeping only skill edits that better predict users' next actions. The method needs no live test environment, because next-action accuracy on held-out logs tracked live success. Unlike earlier methods such as Agent Workflow Memory, which need goal-labeled examples or a live environment, TeleTune guesses each session's goal and uses wrong guesses to suggest edits to a text skill library.

    Image from @rohanpaul_ai's post
  11. LeiphoneNewsAI score42

    Credo moves into optical chips with DSP, PIC and diagnostics in one 1.6T module

    AICredo's latest full-DSP module, Cardinal 802, uses a 4×200G design aimed at both 800G and 1.6T, after the company expanded its ZeroFlap optical module line from 800G to 1.6T over the past year. The company also added Kfir200 silicon photonics PIC from its DustPhotonics acquisition and a PILOT diagnostics platform to the module.

  12. The DecoderNewsAI score53

    Anthropic's Claude Science builds the first complete ultraviolet map of the sky

    AIAnthropic's Claude Science mapped the entire sky in ultraviolet light for the first time. AI agents downloaded data from multiple space missions, calibrated it, merged it, and used inpainting to fill gaps left by NASA's GALEX mission. In tests, predictions averaged about a ten percent deviation from actual measurements.

Oct 8

Oct 8Thu
  1. PandailyNewsAI score46

    ByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression

    AIByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.

  2. PandailyNewsAI score52

    Huawei's KV cache storage faces a missing SSD endurance standard

    AIHuawei's OceanStor M900 and Nvidia's CMX move reusable inference KV cache into a shared storage tier, but no agreed SSD specification exists for it. A storage executive said endurance requirements for one design rose from 3 to 9 drive writes per day and could change again. Industry sources expect convergence to take 6 to 12 months, with another year for development and validation.

  3. SiliconANGLE · AINewsAI score34

    Caterpillar and CoreWeave Shorten Physical AI Learning Loop for Construction Machines

    AICaterpillar and CoreWeave are working to cut the time needed to train construction machines from months to hours, according to Caterpillar's Brandon Hootman. Caterpillar's ecosystem holds about 18 petabytes of federated machine, dealer and customer data, and a single machine can generate terabytes of LiDAR, camera and control data in a day. The partners use AI models, working with Nvidia, to annotate incoming field data so it can feed simulation and training within the same workday.

  4. SiliconANGLE · AINewsAI score27

    Kore.ai launches Autoloop to keep tuning enterprise AI agents after deployment

    AIKore.ai launches Autoloop, an optimization engine that keeps adjusting the AI agents customers build on its Kore.ai Agent Platform to hit goals such as task completion, business-rule adherence and cost, including after deployment. Autoloop writes the first version of each agent from a company's operating procedures, and uses Kore.ai's StateTrace layer to flag where production agents miss goals. The engine is available now to all customers on the Artemis edition of the Kore.ai Agent Platform.

  5. SiliconANGLE · AINewsAI score62

    OpenAI publishes 722 AI-generated math papers on GitHub

    AIOpenAI has published 722 math papers generated by an unreleased AI model, posted to GitHub, spanning about 20 mathematical subfields. The article reports partial progress on the Riemann hypothesis, proving the quasi-Riemann hypothesis without a complete answer, and new results on matrix multiplication, integer multiplication, and Navier–Stokes equations. OpenAI plans to release Lean proofs for more of the papers and to fund review programs for AI-generated math.

  6. Databricks BlogOfficialAI score40

    Databricks Launches Beta Workday Data Connect Federation for Unity Catalog

    AIDatabricks has released a public Beta of a Workday Data Connect federation connector for Unity Catalog, letting teams query Workday HR and finance data in place without copying it. Workday administrators share approved tables through Workday Data Cloud, and Databricks administrators create an OAuth connection and foreign catalog to govern access. The data is read-only, and Workday remains the system of record.

  7. Databricks BlogOfficialAI score40

    Funke Brings Native HL7v2 Parsing to Databricks Lakehouse

    AIDatabricks has released Funke, a Python and PySpark library and deployable pipeline that parses HL7v2 healthcare messages into native Spark types while preserving the full message hierarchy. It succeeds Smolder, the Scala data source Databricks open-sourced in 2021, and ingests through Auto Loader into Unity Catalog bronze and silver tables. Users can query segments, fields, components, and subcomponents directly with DataFrame or Spark SQL expressions.

  8. Databricks BlogOfficialAI score38

    Lakebase Branches Give Parallel Coding Agents Isolated Databases

    AIDatabricks introduces database branching in Lakebase Postgres, letting each coding agent work in its own isolated database branch created in under a second regardless of size. Branches use copy-on-write storage, consuming extra space only as they diverge, and scale to zero when idle so unused branches incur no compute cost. Schema changes are tracked in code and promoted to the parent branch through migrations rather than merged back, and ephemeral branches are created per pull request for testing.

  9. SiliconANGLE · AINewsAI score26

    CoreWeave launches Forge to speed up the AI improvement loop

    AICoreWeave Forge is a platform that runs, observes, curates, improves, and evaluates AI agents and models in a repeating loop, according to CoreWeave vice president of product marketing Susanne Seitinger. The platform adds Agent Lens, Registry, RL Rollouts, model distillation, and the generally available AI Research and Iteration Agent, or ARIA. CoreWeave also introduced a partner network of tested, co-engineered integrations.

  10. Databricks BlogOfficialAI score29

    Biomedical Imaging's Real Bottleneck Is Data Access, Not AI Models

    AIHospitals, academic centers, medtech firms, and pharma companies all face the same obstacle: imaging data is locked in clinical systems and hard to share. The EXAM study across 20 institutions showed federated learning, which shares model weights rather than patient data, improved AUC by 16% on average. Collaboration remains difficult due to scanner and protocol heterogeneity, privacy governance, and the lack of a common data substrate.

  11. AnthropicOfficialAI score57

    Astrophysicist uses Claude to build first complete ultraviolet sky map

    AIAn astrophysicist worked with Claude Science to create the first complete ultraviolet map of the sky, covering regions never observed in UV. Claude located existing datasets, combined them, and filled gaps with statistical inference, taking a few days rather than weeks of human work. The map is presented as a teaching tool and an example of low-priority scientific work that AI now makes feasible.

  12. GoodfireOfficialAI score22

    Alzheimer's Translation Challenge opens registration for spring 2027 competition

    AIThe Alzheimer's Translation Challenge, led by Primamente and AlzData with NVIDIA, Hugging Face, Nebius, Talisman Therapeutics, Ultima Genomics, Cellanome, Prime Intellect, and Boltz, is now open for registration. The competition starts in spring 2027, according to the post.

  13. GoodfireOfficialAI score25

    Goodfire launches a challenge to build AI models on a dataset

    AIParticipants will use a dataset to build AI models, evaluated through a series of evals ranging from general benchmarks to more complex tasks. Top teams will be shortlisted and have their experimental hypotheses tested in Prima Mente's wet lab.

  14. GoodfireOfficialAI score44

    Alzheimer's Translation Challenge Built on 150M-Cell Atlas

    AIThe Alzheimer's Translation Challenge is built on a new atlas of 150M cells, covering neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations with multi-modal readouts. The data will be made available through the AD workbench and Prima Mente's modeling platform.

  15. NVIDIA Technical BlogOfficialAI score29

    NVIDIA KGMON Places Second in KDD Cup 2026 Data Agents Competition

    AIThe NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a smaller, clearer, and easier-to-verify agent harness. The competition required agents to answer natural-language questions over heterogeneous sources, including databases, CSV and JSON files, prose documents, PDFs, and briefing videos.

  16. DatabricksOfficialAI score32

    Databricks' Vibe Data Modeling builds business-specific data models with an agent

    AIDatabricks introduced Vibe Data Modeling, an open-source agent that helps teams build, validate, and evolve business-specific data models. It applies roughly 250 modeling rules while keeping data modelers and business stakeholders involved. Teams can start from 40 industry models as a baseline and iterate toward models that reflect how their business operates.

    Video from @databricks's post
  17. Google Cloud TechOfficialAI score40

    Google Cloud's borderless Lakehouse lets Gemini query multicloud data directly

    AIGoogle Cloud's borderless Lakehouse lets Gemini query data on AWS and Azure without variable egress fees. It reads directly from Salesforce Data 360, SAP, ServiceNow, and Workday without copying data. It also federates open Apache Iceberg tables across Databricks Unity, Snowflake Horizon, and AWS Glue.

    Video from @GoogleCloudTech's post