Skip to content

#Data/Training

Oct 8

TodayOct 8Thu80 items
  1. Pandaily55

    Chinese Team Publishes 3D Cell Atlas of Rice's Full Life Cycle in Cell

    A Chinese-led team published in Cell a three-dimensional spatiotemporal cell atlas covering rice from germinating seed to grain fill, along with a public portal and the RICE scGPT single-cell foundation model. The atlas combines single-nucleus RNA sequencing with BGI's Stereo-seq spatial transcriptomics across 10 organ and tissue types and 61 stages, defining 119 cell types and 133 subtypes.

  2. 雷峰网 Leiphone46

    Huawei and Nvidia Move KV Cache Into Dedicated Storage, but SSD Specs Remain Undefined

    Huawei launched the OceanStor M900 and Nvidia introduced CMX Context Memory Storage to manage inference KV Cache in tiered storage, reducing repeated computation on cache hits. Sources say SSD requirements for this workload are not standardized: some plans have moved from 3 DWPD to 9 DWPD, and the industry may need 6–12 months to reach a common product standard.

  3. 雷峰网 Leiphone54

    Simplexity Robotics uses robots to tend CNC lathes at 0.5 mm tolerance

    At IROS 2026, Simplexity Robotics presented a robot that autonomously tends two CNC machines, grasping workpieces, loading them into chucks, and placing finished parts, with 0.5 mm clearance. The system was trained on only 600 real robot trajectories, using the SimpleWAM world action model, DRAM recurrent memory, and DPE candidate action scoring. The talk was reported as an edited transcript from the conference presentation.

  4. 雷峰网 Leiphone38

    BASAL Intelligence, founded by Tsinghua AIR PhD graduate Li Jianxiong, bets on action-native embodied models

    BASAL Intelligence (本溯智能), founded by Tsinghua AIR's first PhD graduate Li Jianxiong, has raised funding led by Fifth Source Capital, with Ivy Capital, Linear Capital and WestSummit Capital participating. The company is pursuing an action-native, in-context embodied foundation model that lets robots adapt to new scenes in minutes from a few human demonstrations and self-attempts, rather than retraining.

  5. MarkTechPost48

    Laya Open-Source Decision Engine Tutorial: Zero-Shot Decisions and Calibration

    Laya is a 421-million-parameter non-autoregressive decision engine from Convai Innovations that returns calibrated option probabilities in a single forward pass with zero output tokens. This tutorial tests its zero-shot accuracy, probability calibration, temperature fitting, and abstention gating on the CLINC150 banking intent dataset.

  6. SiliconANGLE · AI34

    Caterpillar and CoreWeave Shorten Physical AI Learning Loop for Construction Machines

    Caterpillar and CoreWeave are working to cut the time needed to train construction machines from months to hours, according to Caterpillar's Brandon Hootman. Caterpillar's ecosystem holds about 18 petabytes of federated machine, dealer and customer data, and a single machine can generate terabytes of LiDAR, camera and control data in a day. The partners use AI models, working with Nvidia, to annotate incoming field data so it can feed simulation and training within the same workday.

  7. The Guardian · AI40

    Kenya's Kakuma refugees power tech microwork for dwindling, uncertain pay

    Refugees in Kenya's Kakuma camp are doing AI-related microwork, such as research for RWS's AOP Connect platform, where pay comes as discretionary "rewards" of up to $500 rather than guaranteed wages. Interviews with more than two dozen refugees found most earn far less than the maximum, and entry-level remote tech work is shrinking. Refugees who are unable to legally work in Kenya say they accept these gigs despite opaque pay criteria.

  8. The Decoder62

    Biohub coordinates $1.8 billion effort to build AI models that predict cell behavior

    Biohub, the nonprofit backed by Mark Zuckerberg and Priscilla Chan, is coordinating a $1.8 billion effort covering data, lab equipment, and compute. Meta, Google DeepMind, and Isomorphic Labs are contributing a combined $300 million, and the US Department of Energy is investing over $500 million over five years. Commercial funders get one year of exclusive access to data they paid for, while government-funded data will be public without that restriction, and a first dataset is expected in about a year.

  9. Databricks Blog40

    Databricks Launches Beta Workday Data Connect Federation for Unity Catalog

    Databricks has released a public Beta of a Workday Data Connect federation connector for Unity Catalog, letting teams query Workday HR and finance data in place without copying it. Workday administrators share approved tables through Workday Data Cloud, and Databricks administrators create an OAuth connection and foreign catalog to govern access. The data is read-only, and Workday remains the system of record.

  10. Databricks Blog40

    Funke Brings Native HL7v2 Parsing to Databricks Lakehouse

    Databricks has released Funke, a Python and PySpark library and deployable pipeline that parses HL7v2 healthcare messages into native Spark types while preserving the full message hierarchy. It succeeds Smolder, the Scala data source Databricks open-sourced in 2021, and ingests through Auto Loader into Unity Catalog bronze and silver tables. Users can query segments, fields, components, and subcomponents directly with DataFrame or Spark SQL expressions.

  11. Databricks Blog38

    Lakebase Branches Give Parallel Coding Agents Isolated Databases

    Databricks introduces database branching in Lakebase Postgres, letting each coding agent work in its own isolated database branch created in under a second regardless of size. Branches use copy-on-write storage, consuming extra space only as they diverge, and scale to zero when idle so unused branches incur no compute cost. Schema changes are tracked in code and promoted to the parent branch through migrations rather than merged back, and ephemeral branches are created per pull request for testing.

  12. Prime Intellect52

    Alzheimer's Translation Challenge launches with 150M cell atlas for AI hypothesis discovery

    Prima Mente and AlzData are launching the Alzheimer's Translation Challenge, a global AI competition to discover new therapeutic hypotheses for Alzheimer's disease. The challenge centers on a 150M cell atlas of neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations, with multi-modal readouts. Top teams will have their hypotheses tested in Prima Mente's wet lab, and the data will be available through the AD workbench, Hugging Face, and Prima Mente's modeling platform.

  13. SiliconANGLE · AI36

    CoreWeave Forge Aims to Speed Up the AI Agent Improvement Loop

    CoreWeave launched Forge, a platform meant to help teams run, observe, evaluate and improve AI agents continuously through a five-stage loop. The offering includes Agent Lens, Registry, RL Rollouts, model distillation, and the AI Research and Iteration Agent (ARIA), which is now generally available. CoreWeave also introduced a partner network of tested, co-engineered integrations, according to Susanne Seitinger, vice president of product marketing.

  14. Databricks Blog29

    Biomedical Imaging's Real Bottleneck Is Data Access, Not AI Models

    Hospitals, academic centers, medtech firms, and pharma companies all face the same obstacle: imaging data is locked in clinical systems and hard to share. The EXAM study across 20 institutions showed federated learning, which shares model weights rather than patient data, improved AUC by 16% on average. Collaboration remains difficult due to scanner and protocol heterogeneity, privacy governance, and the lack of a common data substrate.

  15. Sara Hooker10

    It also isn't domain specific, which is important for @adaption_ai since many of the datasets we work with are from domains that are not verifiable. We cover 44 domains which radically differ and have their own nuances.

    It also isn't domain specific, which is important for @adaption_ai since many of the datasets we work with are from domains that are not verifiable. We cover 44 domains which radically differ and have their own nuances.

  16. Sara Hooker17

    @adaption_ai discovery agent reviews past AutoScientist training runs and identify critiques. These end up being at the edge of current capabilities and domain specific. We use that checklist to filter training data automatically.

    @adaption_ai discovery agent reviews past AutoScientist training runs and identify critiques. These end up being at the edge of current capabilities and domain specific. We use that checklist to filter training data automatically.

  17. OpenRouter29

    Mercury Decide from @_inception_ai is now available with zero data retention on OpenRouter The fastest growing decision model on OpenRouter last week now has a paid ZDR endpoint, alongside the free one $0.02/M input (50% off the $0.04 list price). Output and cached input are free. 66K context http://openrouter.ai/inception/mercury-decide

    Mercury Decide from @_inception_ai is now available with zero data retention on OpenRouter The fastest growing decision model on OpenRouter last week now has a paid ZDR endpoint, alongside the free one $0.02/M input (50% off the $0.04 list price). Output and cached input are free. 66K context http://openrouter.ai/inception/mercury-decide

  18. Anthropic57

    Astrophysicist uses Claude to build first complete ultraviolet sky map

    An astrophysicist worked with Claude Science to create the first complete ultraviolet map of the sky, covering regions never observed in UV. Claude located existing datasets, combined them, and filled gaps with statistical inference, taking a few days rather than weeks of human work. The map is presented as a teaching tool and an example of low-priority scientific work that AI now makes feasible.

  19. wh22

    Today, we are sharing more about our post-training infrastructure @ProximalHQ We build on top of @modal GPUs and Kubernetes to support our experiments. Had a ton of fun working on this, and we tried to make the blog as educational as possible! More cool details below :)

    Today, we are sharing more about our post-training infrastructure @ProximalHQ We build on top of @modal GPUs and Kubernetes to support our experiments. Had a ton of fun working on this, and we tried to make the blog as educational as possible! More cool details below :)

  20. Epoch AI · The Epoch Brief49

    Epoch AI's October 2026 Brief Covers AI Agents, Falling Costs, and China's Chip Exposure

    Epoch AI estimates the AI chips shipped through 2027 could support about 30 to 170 million concurrent frontier-model agents, or nearly 2 billion with cheaper models. Its researchers find the cost of a fixed level of AI performance has fallen about 47% per quarter over the past three years. The newsletter also reports China's semiconductor supply-chain exposure is 2.7 times that of the US.

  21. Google Research20

    Interested in privacy-preserving synthetic data? Catch Alex Bie at the @COLM_conf Google booth (#107) today, at 2:00 PM PT for an encore presentation of ContinuousBench, evaluating knowledge transfer and data contamination in DP synthesis.@GoogleDeepMind Join the conversation at #COLM2026!

    Interested in privacy-preserving synthetic data? Catch Alex Bie at the @COLM_conf Google booth (#107) today, at 2:00 PM PT for an encore presentation of ContinuousBench, evaluating knowledge transfer and data contamination in DP synthesis.@GoogleDeepMind Join the conversation at #COLM2026!

  22. The Decoder65

    Anthropic launches Claude Dashboards and Motion features in beta

    Anthropic launched two beta features for Claude: Dashboards, which turns connected data sources like BigQuery, Databricks, Snowflake, or Salesforce into auto-updating live dashboards from text prompts, and Motion, which creates animated explainer videos from text, diagrams, and images. Dashboards is available to paid users and Motion to Team and Enterprise plans, while Docs, Slides, and Design leave beta and work across all plans, including free accounts.

  23. Goodfire22

    The Alzheimer’s Translation Challenge is led by @primamente and @AlzData, with @nvidia, @huggingface, @nebiusai, Talisman Therapeutics, @ultimagenomics, @Cellanome, @PrimeIntellect, and @boltz_bio. Registration is open now. Competition starts spring 2027: https://primamente.com/challenges

    The Alzheimer’s Translation Challenge is led by @primamente and @AlzData, with @nvidia, @huggingface, @nebiusai, Talisman Therapeutics, @ultimagenomics, @Cellanome, @PrimeIntellect, and @boltz_bio. Registration is open now. Competition starts spring 2027: https://primamente.com/challenges

  24. Goodfire25

    Participants will use the dataset to build AI models, assessed through a series of evals from general benchmarks to more complex tasks. Top teams will be shortlisted, and their experimental hypotheses tested in Prima Mente's wet lab.

    Participants will use the dataset to build AI models, assessed through a series of evals from general benchmarks to more complex tasks. Top teams will be shortlisted, and their experimental hypotheses tested in Prima Mente's wet lab.

  25. Goodfire44

    The Alzheimer’s Translation Challenge is built on a new 150M-cell atlas of neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations, with multi-modal readouts. The data will be made available through the AD workbench and Prima Mente's modeling platform.

    The Alzheimer’s Translation Challenge is built on a new 150M-cell atlas of neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations, with multi-modal readouts. The data will be made available through the AD workbench and Prima Mente's modeling platform.

  26. Tessl Blog29

    One Brain Means Owning Your Organizational Memory

    Leapfrog, a small team doing high-volume AI visual and production work for fashion and brand clients, is building a "one brain" system that makes company knowledge and client context searchable through natural-language agents. The starter stack described is OpenClaw in a sandbox, a GitHub repository, Obsidian on the local machine, and Telegram as the access point. The system's research structure had roughly 1,200 files at the time of the talk.

  27. Databricks32

    Databricks' Vibe Data Modeling builds business-specific data models with an agent

    Databricks introduced Vibe Data Modeling, an open-source agent that helps teams build, validate, and evolve business-specific data models. It applies roughly 250 modeling rules while keeping data modelers and business stakeholders involved. Teams can start from 40 industry models as a baseline and iterate toward models that reflect how their business operates.

  28. TechCrunch · AI36

    Ben Affleck's AI expertise goes viral as he explains neural networks and fine-tuning

    Actor Ben Affleck drew attention this week for explaining machine learning concepts, including convolutional neural networks, tensors, and transformers, in several recent interviews. He said he fine-tuned open video models by unfreezing weights and training only the last cinematic layer, using a dataset he built over about eight months for his startup. Affleck said he worries about students and learned helplessness more than Skynet, and predicted AI will be additive to the movie business.

  29. Google Cloud Tech40

    Our borderless Lakehouse lets Gemini query AWS and Azure with no variable egress fees, read directly from Salesforce Data 360, SAP, ServiceNow, and Workday without copying data, and federate open Apache Iceberg tables across Databricks Unity, Snowflake Horizon, and AWS Glue → https://goo.gle/3TT9y4X

    Our borderless Lakehouse lets Gemini query AWS and Azure with no variable egress fees, read directly from Salesforce Data 360, SAP, ServiceNow, and Workday without copying data, and federate open Apache Iceberg tables across Databricks Unity, Snowflake Horizon, and AWS Glue → https://goo.gle/3TT9y4X

  30. Google Cloud Tech12

    Smart Storage is a game-changer for dark, unstructured data. Smart Storage enriches files in place and writes metadata directly onto the source object, so Gemini gets instant context and your security ACLs never drift.

    Smart Storage is a game-changer for dark, unstructured data. Smart Storage enriches files in place and writes metadata directly onto the source object, so Gemini gets instant context and your security ACLs never drift.

  31. Jerry Liu22

    LlamaIndex argues Markdown is the universal format for agents

    LlamaIndex says Markdown has become a universal representation between humans and agents, preserving headings, lists, and tables while remaining readable to models. Since most unstructured documents are not natively in Markdown, the main challenge is the translation layer, which the company addresses with models that convert document containers into Markdown. The quoted post adds that Markdown keeps table columns intact, with HTML used for tables with merged headers.

  32. Jerry Liu41

    OpenDocRouter is basically OpenRouter for document OCR You can compare and use all your favorite OCR models under a single API / billing interface. They range from some of the lightweight/OSS models (e.g. MinerU) to the latest frontier VLMs (Opus 5.5) Come check it out! https://www.opendocrouter.ai/

    OpenDocRouter is basically OpenRouter for document OCR You can compare and use all your favorite OCR models under a single API / billing interface. They range from some of the lightweight/OSS models (e.g. MinerU) to the latest frontier VLMs (Opus 5.5) Come check it out! https://www.opendocrouter.ai/