Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 30

Sep 30Wed
  1. FireworksOfficialAI score34

    GLM 5.3 Flash now available for training on Fireworks' Serverless API

    AIFireworks AI has made GLM 5.3 Flash available for training through its Serverless Training API, open to all users. The model supports both vision and text inputs. Fireworks says it performs well on its benchmarks for agentic coding, document analysis, and tool use while remaining cost-efficient to serve.

  2. Google Cloud TechOfficialAI score28

    Agent Clinic Ep 3 builds automated eval suite for LangGraph agent

    AITerminal test runs miss multi-turn agent regressions, so Agent Clinic Episode 3 builds an automated eval suite for a LangGraph agent in 60 minutes. The post presents a four-step framework for moving from informal checks to benchmarking AI agents, with a link to the full guide.

    Image from @GoogleCloudTech's post
  3. Google Cloud TechOfficialAI score20

    Google Cloud Model Garden adds scale-to-zero to power down idle GPUs

    AIGoogle Cloud Model Garden now lets users enable scale-to-zero to automatically shut down GPU instances when no incoming requests are active. The feature is presented as a way to avoid paying for idle GPU capacity.

  4. TypeSafe AIOfficialAI score27

    Jev reranking beats GPT-5 Mini on sales data retrieval

    AIJev reranking retrieves Rox sales data 20x faster, 10x cheaper, and 12% more accurate than GPT-5 Mini. The benchmark compared Jev classification against LLM-based reranking for pulling transcripts, emails, CRM notes, news, and documents.

  5. Higgsfield AI 🧩OfficialAI score37

    Higgsfield adds computer use via ChatGPT extension inside Codex

    AIHiggsfield's full interface is now available inside Codex, powered by GPT-6.1 Sol, with access to local files and automation skills. The integration lets users automate creative workflows end-to-end within ChatGPT.

    Video from @higgsfield's post
  6. GitHubOfficialAI score42

    Project HydraFusion Now Available in GitHub Copilot App and VS Code

    AIProject HydraFusion is now available in the GitHub Copilot app and VS Code, where users can select it like any other model. Behind the scenes, it routes a task across multiple models to draft, critique, revise, or escalate, then returns one combined result.

    Video from @github's post
  7. Nathan LambertXAI score47

    “Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coor...

    AI“Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service” It’s the API company’s problem if their model can be manipulated like this. Add KYC

  8. NVIDIA AIOfficialAI score40

    NVIDIA Shows Visual AI Agent Built in Under 30 Minutes

    AINVIDIA says a single prompt can build and deploy a visual AI agent for a manufacturing line in under 30 minutes, with alerts, video search, and incident reports. The method uses the new Build Vision AI skill in NVIDIA VSS Blueprint 3.3, and a tutorial is available for readers who want to build one.

    Video from @NVIDIAAI's post
  9. Demis HassabisXAI score62

    Google DeepMind's SynthID Bio watermarks AI-designed proteins in Nature study

    AIGoogle DeepMind reports that AI-designed proteins can be synthesized and watermarked using its new SynthID Bio method, published in Nature. The team says the work is a step toward biosecurity in AI-driven biology and is open-sourcing the SynthID Bio tools for the research community.

    Why it matters: The source reports a published Nature study and open-sourced tools, showing a concrete method for watermarking AI-designed proteins against misuse.

  10. Jensen HuangXAI score62

    Industry leaders sign White House Accord on Super Intelligence safety commitments

    AIJensen Huang says leaders across the industry signed the White House Accord on Super Intelligence at the White House. The accompanying document says each company should run internal controls, an independent external auditor, and board-level oversight for its frontier models.

    Why it matters: The source gives the accord's four layers of controls and audits, showing how the signing parties plan to verify frontier model safety in practice.

    Image from @JensenHuang's post
  11. FireworksOfficialAI score30

    Fireworks' Ember-1 matches Kimi K3 on Vals with fewer reasoning tokens

    AIFireworks' Ember-1 performs close to Kimi K3 on Vals' finance and legal benchmarks while using fewer reasoning tokens per turn. Fewer tokens per turn lower cost and speed up agent loops, and Ember-1 is available on Fireworks Serverless.

  12. Google WorkspaceOfficialAI score22

    Bci reaches 90% AI adoption, saving over 7,000 hours with Gemini

    AIJuan Burgueño says Bci has reached over 90% AI adoption, saving more than 7,000 hours by using Google Workspace and Gemini. The bank has scaled to over 2,000 custom Gems to accelerate innovation and build an AI-first bank.

    Video from @GoogleWorkspace's post
  13. FireworksOfficialAI score61

    Fireworks adds GLOBAL multi-region deployments under one endpoint

    AIFireworks has added a GLOBAL option that lets one deployment run across regions behind a single endpoint and identity. The scheduler draws compatible capacity from the broadest allowed pool while respecting hardware, quota, reliability, and data-residency constraints. In a seven-day observational study, multi-region deployments showed a 99.992% request success rate versus 99.269% for single-region deployments, though the authors say this is not causal.

    Why it matters: The article explains how one deployment draws capacity across regions, showing how scheduling and routing changed to remove a single-region capacity ceiling.

  14. Ant LingOfficialAI score22

    Ling-3.1-flash scores 65.35 on HealthBench Professional benchmark

    AIAnt Ling reports that Ling-3.1-flash scored 65.35 on HealthBench Professional, a healthcare evaluation rather than clinical certification. The model was trained with feedback from AQ and Haodf.com (Good Doctor Online) healthcare scenarios.

    Image from @AntLingAGI's post
  15. Ant LingOfficialAI score31

    Ant Ling model turns plain-language prompts into interactive Three.js pages

    AIAnt Ling can convert plain-language prompts into standalone, interactive Three.js pages covering topics such as an internal combustion engine, an optical-disc reader, paramecium organelles, and vector-field divergence. The output is runnable code rather than just an explanation.

    Video from @AntLingAGI's post
  16. Ant LingOfficialAI score38

    Ling-3.1-flash ports C image library to Rust with 8.015× speedup

    AIAnt Ling reports that its Ling-3.1-flash model completed a roughly 20-hour Rust port of a C image library. After a performance regression caused by busy-waiting workers and a parallelism adjustment, the model recovered and reached an 8.015× speedup. All 30 correctness checks passed.

    Image from @AntLingAGI's post
  17. Ant LingOfficialAI score28

    Ant Group's Tiger Agent runs Ling-3.1-flash for desktop task automation

    AIAnt Group's internal Tiger Agent uses Ling-3.1-flash to plan tasks, while its desktop agent provides browser, file, terminal, and live preview capabilities. A GitHub Trending demo shows the system turning research and semantic grouping into a concise brief.

    Video from @AntLingAGI's post
  18. Ant LingOfficialAI score46

    Ant Ling releases Ling-3.1-flash with 1M-token context, plans open-source

    AIAnt Ling introduced Ling-3.1-flash, a model with about 560B total parameters, about 25B active per token, and up to a 1M-token context window. The company plans to open-source the model soon. It reports 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional across work, coding, and healthcare tasks.

    Image from @AntLingAGI's post
  19. NVIDIA AIOfficialAI score27

    NVIDIA NeMo Relay Traces Hermes Agent Runs in Arize Phoenix

    AINVIDIA and Nous Research published a hands-on walkthrough of NVIDIA NeMo Relay for collecting traces from Hermes Agent. The guide runs two example scenarios and shows the agent's calls and retries in Arize Phoenix. It also covers how Nous used traces and task results to evaluate fixes across repeated runs.

    Video from @NVIDIAAI's post
  20. DeepSeek HarnessXAI score62

    DeepSeek Harness v0.2 preview launches as a desktop app for macOS and Windows

    AIDeepSeek releases the DeepSeek Harness v0.2 preview with a desktop app for macOS and Windows. The release adds a plugin manager for installing, disabling, and uninstalling plugins without terminal commands, plus an experimental creator mode that generates plugins from user descriptions. The company says DeepSeek Harness is now the most widely used coding agent among users of the official DeepSeek API by DAU and daily sessions.

  21. Google GeminiOfficialAI score60

    Gemini skills roll out globally and expand to Google Workspace customers

    AISkills are rolling out globally in Gemini today. They will expand to Google Workspace business, enterprise, nonprofit, and education customers in the coming weeks. The post links to a blog for how users can use skills to handle repetitive tasks.

    Why it matters: The post states the rollout scope and timing for Gemini skills, which matters to Workspace admins and customers planning their own adoption.

  22. Google GeminiOfficialAI score45

    Gemini skills now proactive, stackable, and support reference files

    AIGemini can build custom skills from chats and apply a saved skill automatically when a prompt matches it. Multiple skills can be stacked for larger tasks, such as combining a personal writing style skill with a brand guidelines skill. Starting today, skills can include reference files such as plain text documents, PDFs, or images, with sharing and Google Drive file support coming soon.

  23. Google GeminiOfficialAI score34

    Gemini Gems will migrate automatically into Skills

    AIGoogle says Gemini Gems will be replaced by Skills as the tool for tailoring instructions to specific tasks, starting November for personal accounts. Workspace business, enterprise, and nonprofit customers transition in March 2027, and Workspace education customers in June 2027. Existing Gems will migrate into Skills automatically, with transition details in Google's help center.

  24. Microsoft ResearchOfficialAI score29

    Microsoft Research ML system predicts space-weather damage 30-60 minutes early

    AIMicrosoft Research has developed a machine learning system that predicts where extreme space-weather events are likely to damage power systems 30-60 minutes before a storm arrives. Such storms can also degrade GPS accuracy and satellite operations, so advance warning could help operators prepare.

    Video from @MSFTResearch's post
  25. Google · Gemini appOfficialAI score42

    Google Gemini Adds Reusable Skills to Replace Gems Over Coming Months

    AIGoogle is rolling out skills in Gemini chat globally, letting users save frequently used instructions and invoke them by typing a forward slash and the skill name. Skills will replace Gems, which Google will remove starting in November for personal accounts, March 2027 for Workspace business, enterprise and nonprofit customers, and June 2027 for Workspace education customers. Gems will be automatically migrated into skills.

  26. Google Cloud · AI & Machine LearningOfficialAI score45

    Google Cloud Launches Preview of CLI Remote MCP Server for AI Agents

    AIGoogle Cloud has introduced the Google Cloud CLI remote MCP server in preview, giving AI agents access to gcloud and bq command-line operations through two tools, run_gcloud_command and run_bq_command. The server runs in an isolated, network-restricted execution sandbox on Google Cloud, so teams need no local CLI binaries, and calls are authenticated through Agent Identity, OAuth 2.0, and IAM, with Model Armor screening and Audit Logs available.

  27. Microsoft ResearchOfficialAI score46

    Machine learning system forecasts space-weather grid risk for 66,935 U.S. substations

    AIMicrosoft Research intern-developed machine learning pipeline forecasts location-specific geomagnetic risk for 66,935 substations in the continental United States. It combines solar-wind observations, AE and Dst forecasts, geological conductivity and grid data to estimate risk 30 to 60 minutes ahead. The pipeline detected nearly 80% of major space-weather events during the evaluation period.