Skip to contentSkip to stories

Updated

#Data/Training

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 10

TodayOct 10Sat
  1. PandailyNewsAI score44

    Maniformer opens a crowdsourcing platform that pays people to record tasks for robots

    AIShanghai company Maniformer has opened a crowdsourcing platform that pays ordinary people to record everyday tasks, such as folding clothes or making tea, as training data for robots. The platform lists 22 categories, over 5,000 task types and more than 50,000 real scenes, and nearly 20,000 users signed up during a closed test from August 15 to September 15. Maniformer also announced a RMB 100 million subsidy plan and a goal of 1 million daily active users and 1 million robot trainers within two years.

  2. LeiphoneNewsAI score31

    Metabee launches Mifeng Pai, a global crowdsourcing platform for physical AI data

    AIMetabee launched Mifeng Pai, which it calls the world's first full-category crowdsourcing platform for high-quality physical AI data, at a Shanghai event on September 23. Contributors earn from tasks such as folding clothes or making tea, with payouts within two minutes, and the platform covers 22 categories, more than 5,000 task types and 50,000 real-world scenes. Metabee also announced a 100 million yuan subsidy plan and a robot trainer certification program.

Oct 9

Oct 9Fri
  1. vLLM BlogOfficialAI score62

    vLLM adds support for NVIDIA Vera Rubin NVL72 with 7.8x throughput over GB200

    AIvLLM now supports NVIDIA Vera Rubin NVL72, with daily container builds and support for models from DeepSeek, Moonshot AI, Z.ai, and MiniMax. In early AgentX benchmarks, vLLM running MiniMax M3 delivered up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity. The post is an early look, and the team expects further gains from ongoing optimizations.

  2. RadixArkOfficialAI score28

    Miles v0.1.2 adds Kubernetes support and torchtitan training backend

    AIRadixArk releases Miles v0.1.2, adding score centering for stable async RL and an experimental native Kubernetes backend. RL jobs can run as ordinary cluster workloads, and orchestration can restart while training continues. The release also adds torchtitan as a third training backend alongside Megatron and FSDP, plus support for DeepSeek-V4.1-Flash and MiMo-V2.6-Flash-RL.

    Image from @radixark's post
  3. Rohan PaulXAI score46

    Sabi raises $50M seed round for a brain-signal baseball cap

    AISabi has raised a $50M seed round to build a baseball cap that reads brain signals to generate prompts for AI models. Sabi says the cap can hold up to 100,000 sensors and predict a user's next three to four keystrokes before they type them. The round was led by Vinod Khosla with Accel, DST Global, Initialized Capital and Kevin Weil, at a reported $600M valuation.

    Image from @rohanpaul_ai's post
  4. SantiagoXAI score43

    Sabi raises $50 million to build a wearable brain-signal-to-text cap

    AISabi has raised $50 million from Khosla Ventures, Accel, Initialized, Kevin Weil and DST Global to build a wearable brain-computer interface cap. The author says the company already has a TSMC-fabricated chip that reads brain signals without touching the scalp, with custom sensors and an in-house AI model that decodes the signals into text. The author notes that the device is a wearable cap, not an implant.

  5. LeiphoneNewsAI score42

    Tian Keyu's 10-person world model lab reaches a $200M valuation after ByteDance dispute

    AIA 10-person stealth lab founded by Peking University PhD student Tian Keyu raised $30 million from Wolves VC and IDG at a $200 million valuation, according to Leiphone. Tian, who was sued by ByteDance for 8 million yuan over a 2024 sabotage incident, plans a base model trained on about 100 million hours of video using a 200,000-symbol visual vocabulary, targeting a 2027 release.

  6. GeekParkNewsAI score62

    Krinwave raises 400 million yuan to bring brain imaging ultrasound to AI

    AIGeekPark reports that Krinwave, a Shenzhen company, has completed a new 400 million yuan round with XVC and Sequoia China as new investors. The company says its low-frequency ultrasound captures brain structures through the skull, which conventional high-frequency devices cannot, and it aims to supply hardware and its kOS software platform to brain-computer interface and NeuroAI groups.

  7. QbitAINewsAI score44

    Sharpa unveils D01 humanoid, W02 hand and AE01 glove at IROS

    AISharpa has unveiled the D01 humanoid robot, the W02 dexterous hand and the AE01 exoskeleton data glove at IROS. The article says the D01 has full-body electronic skin, and the W02 has 21 active degrees of freedom and a hand about 30% smaller than the W01.

Oct 8

Oct 8Thu
  1. LeiphoneNewsAI score38

    BASAL Intelligence, founded by Tsinghua AIR's Li Jianxiong, raises funding for action-native embodied models

    AIBASAL Intelligence, founded by Tsinghua AIR's first PhD graduate Li Jianxiong, has secured lead funding from Five Source Capital, with Ivy Capital, Linear Capital and WestSummit Capital participating. The company says it is betting on an action-native In-Context embodied foundation model that lets robots learn on site from a few human demonstrations and trial attempts within minutes, without retraining after a scene change.

  2. Prime IntellectOfficialAI score52

    Alzheimer's Translation Challenge launches with 150M cell atlas for AI hypothesis discovery

    AIPrima Mente and AlzData are launching the Alzheimer's Translation Challenge, a global AI competition to discover new therapeutic hypotheses for Alzheimer's disease. The challenge centers on a 150M cell atlas of neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations, with multi-modal readouts. Top teams will have their hypotheses tested in Prima Mente's wet lab, and the data will be available through the AD workbench, Hugging Face, and Prima Mente's modeling platform.

    Video from @PrimeIntellect's post
  3. OpenRouterOfficialAI score29

    Mercury Decide now available with zero data retention on OpenRouter

    AIInception's Mercury Decide, the fastest-growing decision model on OpenRouter last week, now has a paid zero data retention (ZDR) endpoint alongside the free one. It is priced at $0.02/M input tokens, half the $0.04 list price, with output and cached input free and a 66K context window.

    Video from @OpenRouter's post
  4. Jerry LiuXAI score41

    LlamaIndex launches OpenDocRouter, a unified API for document OCR models

    AIJerry Liu introduces OpenDocRouter, a unified API for document parsing that serves frontier and open-weight OCR and vision-language models under one billing interface. The service is offered at cost with a small transaction fee, handles rate limits, adds bounding boxes and layout output, and will add new models after benchmarking them on ParseBench.

    Video from @jerryjliu0's post
  5. elvisXAI score34

    Monsoon ASR dataset cuts Bengali Whisper word error rate to 7.65%

    AIVoice Arena's Monsoon ASR dataset fine-tuned Whisper Medium on Bengali FLEURS, reducing LLM word error rate from 85.27% to 7.65%. The corpus spans 100,000 hours across 50 languages, and Voice Arena says more than 80 organisations have asked to license it since its launch a week ago.

  6. IEEE Spectrum · AINewsAI score46

    Nuclear Plants Adopt AI Tools, Led by Atomic Canyon's NIVA Assistant

    AIAtomic Canyon's Nuclear Industry Virtual Assistant (NIVA), developed with nuclear-industry groups, is now available to the entire U.S. fleet of 94 reactors after pilot testing at Constellation Energy plants. Nuclearn says its products have reached more than 65 U.S. partners, and the article says the industry is turning to AI to help manage regulatory paperwork and a shrinking, aging workforce.

  7. X.PINXAI score40

    Tian Keyu's unnamed startup raises nearly $30M at $200M valuation

    AITian Keyu's unnamed startup has raised nearly $30M from investors including 5Y Capital and IDG at a $200M post-money valuation, according to Bloomberg. The company is developing a 200,000-symbol visual vocabulary for video AI and claims it could cut video-generation costs at least tenfold. A full model release is planned for 2027, and the company has no product yet.

    Image from @thexpin's post

Oct 7

Oct 7Wed
  1. Sara HookerXAI score36

    Adaption Labs launches Invent-a-dataset for generating post-training datasets

    AIAdaption Labs has launched Invent-a-dataset, which creates post-training datasets from a natural language description alone. A technical report compares the Invent API against frontier proprietary and open-weight LLMs used for data generation. Sara Hooker's post links to the launch page and the report's technical details.

  2. Daniel HanXAI score48

    Unsloth shows how to fine-tune LLMs into decision models locally

    AIDaniel Han reposts Unsloth's announcement that LLMs such as Qwen3.5 0.8B can be fine-tuned into decision models with a Clef head and LoRA (r=64) on about 3GB to 4GB of VRAM. The post says downstream accuracy rose from 30–37% to 78%, and that Qwen3.5 0.8B's aggregate accuracy on three decision benchmarks went from 20.7% to 74.3%. The training guide and notebooks are linked from the Unsloth GitHub repo.

    Video from @danielhanchen's post
  3. SantiagoXAI score22

    Model infers derived values from document data, computing yearly costs from monthly figures

    AIA new model extracts values absent from a document by computing them from figures that are present, such as deriving a yearly product cost from a monthly price. Santiago says the video shows examples of inferring complex formulas. The background post describes this as Higher-Order Extraction, which deterministically computes needed numbers from raw page values.

  4. Jerry LiuXAI score47

    Jerry Liu introduces OpenDocRouter, a unified API for document parsing models

    AIJerry Liu announces OpenDocRouter, a unified API for document parsing that serves 10 frontier and open-source models at launch. Users can switch models in one line with the same request and markdown output, and each model is scored on ParseBench for quality and cost. Pricing is per token, with failed pages not charged and costs ranging from $0.86 to $48.82 per 1,000 pages.

    Video from @jerryjliu0's post
  5. LlamaIndex 🦙OfficialAI score47

    LlamaIndex launches OpenDocRouter, a single API for document parsing models

    AILlamaIndex announces OpenDocRouter, a single API that gives access to 10 frontier and open-source document parsing models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro and PaddleOCR-VL-1.6. Users can switch models in one line with the same request and markdown output, and the post says every model is scored on ParseBench for quality and cost. Pricing is per token at $0.86 to $48.82 per 1,000 pages, with failed pages not charged and top-ups starting at $25.

    Video from @llama_index's post

Oct 6

Oct 6Tue
  1. Hacker News · AI (150+ points)BlogAI score39

    Penguin Mail 1.0.5 is an open-source Rust email client for Linux with AI

    AIPenguin Mail 1.0.5 is a free, GPL-3.0-or-later email and calendar app for x86_64 Linux that supports Gmail, Microsoft, IMAP and POP3 accounts. The app includes an optional AI assistant that stays off until a model is chosen and can run locally through LM Studio or Ollama, asking before it sends mail or changes settings.

  2. IThome · AINewsAI score41

    Strata engine runs quantized Qwen3.8-Flash-Next 125B on 12GB VRAM GPUs

    AIDeveloper Niko1221 open-sourced Strata, an engine that runs quantized Qwen3.8-Flash-Next, a 125B-parameter model, on consumer GPUs with at least 12GB of VRAM. It loads the MoE model into RAM and keeps only frequently used experts in VRAM, and uses a lightweight model for speculative decoding. On an RTX 5070 with 64GB RAM, the Q2_0 quantization generates 94 tokens per second, and an RX 9070 XT reaches 60 tokens per second.

Oct 5

Oct 5Mon
  1. Google AIOfficialAI score46

    Living Models pairs Gemma 4 with BOTANIC-1 to rank crop gene mutations in minutes

    AILiving Models paired Gemma 4 with BOTANIC-1, a model trained on plant DNA across 320 species, to identify which mutations alter crop traits. In a melon yield test, the pipeline ranked the target mutation first out of 2,494 candidates in under four minutes. Google AI says the approach could help researchers engineer climate-resilient crops faster than traditional breeding trials.

  2. Elad GilXAI score40

    Era launches free simulated enterprises for testing AI agents

    AIEra, launched by Ofir Ehrlich's team, generates a complete simulated company spanning Salesforce, Slack, Jira, Zendesk, Gong, and Deel, plus cloud databases and storage. Agents interact with it through live MCP and API interfaces, and because Era generated the company, it knows the exact ground truth for testing and benchmarking. The post says the product is live today and free.

Oct 4

Oct 4Sun
  1. B A H A R 🧬XAI score22

    Axis raises $12M seed to scale decentralized robot training data

    AIAxis, a physical AI data startup, announces a $12M seed round led by Hack VC, with backing from Nomad Capital, Pi Network Ventures, 10K Ventures, and Ondo. The company says the funding will expand human-in-the-loop teleoperation and GPU augmentation pipelines. Its Axis Hub lets more than 79,000 contributors complete robot manipulation tasks in the browser using MuJoCo-WASM, without expensive GPUs.

    Image from @Bahar3479's post
  2. Jerry LiuXAI score34

    LlamaIndex launches Extract v2.5 document extraction agents, cutting errors on scanned forms

    AILlamaIndex introduced Extract v2.5, a series of agents tuned for document extraction, including cost-effective, agentic, and agentic plus tiers, available in LlamaParse. The company says the agents reduce error rates by 2x or more compared with frontier models at a small fraction of the price, and they handle handwritten and drawn annotations on scanned documents while grounding values in the source text.

    Video from @jerryjliu0's post

Oct 2

Oct 2Fri
  1. Jerry LiuXAI score34

    LlamaIndex's Extract v2.5 agents reason over tables spanning multiple pages

    AILlamaIndex introduced Extract v2.5, a set of document extraction agents that can reconstruct records split across pages and assemble them with thousands of other cells into structured tabular output. The post says the agents handle real-world documents like insurance claims, regulatory filings, and legal schedules, where a record may start on one page and finish on the next. The accompanying background post claims record-spanning-page accuracy rose from 85.5% to 96.5%, and that the agentic tier outperforms Opus 5.5 and GPT-6 Sol at 30% to 4x lower cost.

    Video from @jerryjliu0's post
  2. Sara HookerXAI score26

    Adaption Labs makes its Invent dataset tool available via API

    AIAdaption Labs has made Invent, its tool for generating AI training datasets from a plain-language description, available through an API. Developers can reportedly produce AI-ready training datasets in minutes with a few lines of code, according to the quoted post. Documentation is available at docs.adaptionlabs.ai.

    Image from @sarahookr's post

Oct 1

Oct 1Thu
  1. Jerry LiuXAI score42

    Jerry Liu introduces Extract v2.5, a set of extraction agents tuned for documents

    AIJerry Liu announces Extract v2.5, a series of agents for document extraction in three tiers: cost-effective, agentic, and agentic plus. He says they outperform Opus 5.5 and GPT-6 Sol at 30% to 4x lower cost, and that accuracy on long lists rose from 86.1% to 95.5%. The release also adds advanced citations with bounding boxes and structural reasoning, and is available on LlamaParse.

    Video from @jerryjliu0's post
  2. Mastra BlogOfficialAI score26

    Mastra Platform Adds VPC-Isolated Postgres Databases for Same-Network Access

    AIMastra platform now lets users attach a VPC-isolated Postgres database to any environment, restricting access to resources on the same network. The database cannot be reached from outside the network, so psql connections from external clients return an error. VPC Postgres joins Turso and Neon as managed database options, with MongoDB and Redis coming soon; it requires mastra@1.32.0 or later.

Sep 30

Sep 30Wed
  1. FireworksOfficialAI score32

    Fulcrum's Echo model trained on Kimi K3 LoRA via Fireworks

    AIFulcrum's Echo was trained using Fireworks' serverless training API on a rank 32 LoRA of Kimi K3. The team's writeup describes a "persona post-training" approach and a total training cost of $4.5k.

Sep 29

Sep 29Tue
  1. Mastra BlogOfficialAI score42

    Mastra adds memory hooks to monitor and transform agent observations

    AIMastra says it now lets developers add lifecycle and transform hooks to its observational memory system. Lifecycle hooks log cycle starts and ends with token usage, while transform hooks can modify observations and reflections before they are stored. The release also includes skillResultRedactor(), which replaces skill file contents in tool results with a placeholder so the agent still remembers the skill.

Sep 28

Sep 28Mon
  1. FireworksOfficialAI score22

    Normal Factory's CAD Arena joins the Specialized Intelligence Index

    AINormal Factory joins the Specialized Intelligence Index with CAD Arena, which tests whether AI agents can turn engineering drawings into accurate, editable CAD parts. The benchmark evaluates agents across five CAD platforms, extending the SII into engineering design.

    Image from @FireworksAI_HQ's post
  2. ModelScopeOfficialAI score43

    Jina-OCR-v1 parses full pages into Markdown at 2.57 pages per second

    AIJina-OCR-v1, a 3.4B-parameter MoE model that activates 570M parameters per token, converts entire document pages into structured Markdown at 2.57 pages per second. It scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, 7.4 points above DeepSeek-OCR on the latter, and delivers the highest throughput among 14 evaluated systems at concurrency 32. The model is released under CC BY-NC 4.0, so commercial use requires permission.

    Image from @ModelScope2022's post