Arena thanks a16z for continued partnership and support
AIArena's account publicly celebrates continuing its work with a16z and the team. A quoted congratulatory post from Robert N. Weber praises the founders' rapid progress in building the company.
Updated
Updated
Showing low-relevance items too. Hide low-relevance items
AIArena's account publicly celebrates continuing its work with a16z and the team. A quoted congratulatory post from Robert N. Weber praises the founders' rapid progress in building the company.
AIThe U.S. Department of Labor said it suspended Microsoft and Adobe from its Permanent Labor Certification program, citing multiple active federal investigations. Labor Secretary Keith Sonderling also said no new applications will be accepted for Cognizant, Infosys, Capgemini, Tata, Wipro and HCL. Microsoft said the vast majority of its U.S. employees are Americans and that 80% of its roughly 6,000 H-1B petitions last fiscal year were to extend or change the status of existing employees.
AIArena is recruiting through its careers page, and @johnnnavent is hiring a designer for its consumer app at arena.ai. The designer role calls for deep consumer experience and work at the frontier of AI.
AIArena has secured a $200M Series B, with Lightspeed doubling down on its investment. The company is also launching the Alignment Index, which measures how closely AI behavior aligns with human values in real-world settings. Arena reports annualized revenue above $100M since its Series A, with millions of people helping evaluate frontier models through real-world use.
AILeapfrog, a small team doing high-volume AI visual and production work for fashion and brand clients, is building a "one brain" system that makes company knowledge and client context searchable through natural-language agents. The starter stack described is OpenClaw in a sandbox, a GitHub repository, Obsidian on the local machine, and Telegram as the access point. The system's research structure had roughly 1,200 files at the time of the talk.
AITessl's talk at AI Native DevCon London argues that agent skills, which can be markdown files with instructions and bundled material, act as supply chain components that can shape agent behavior. The author says reading SKILL.md once is insufficient because risks can sit in supporting files, updates, and workspace trust settings. He identifies the danger as the combination of private context, untrusted content, and external communication, and cites research scanning roughly 4,000 public skills for issues including malware-like behavior.
AIElvis Saravia argues that creative work needs its own agent harness rather than coding-focused tools. He highlights Voyager, which he says is an open harness for video, graphics and games that works with local files and drives apps such as Blender, DaVinci Resolve and Unity. The post says it can run models including Opus, Astra and DeepSeek.
AIGoogle Research is hosting a live demonstration of EmbeddingGemma 2 at its COLM booth #107 today at 12:00pm. The open multimodal model unifies text, image, audio, and video representations, with Sahil Dua available to connect with attendees.
AIMozilla.ai's cq project proposes a shared knowledge layer where AI agents capture lessons from non-obvious fixes as structured knowledge units that other agents can later query. The default setup is local-first, using a local SQLite database so nothing leaves the machine, with an option to connect to a remote team server that adds review.
AIMuse, the AI assistant, read the headers of a phishing email, traced its spoofed domains, and drafted an AWS takedown report in about ten minutes. The post's author, @accuratetlm13, reports this after receiving the spam email, framing it as a warning to phishers.
AIJohn Groetzinger, writing in a personal capacity rather than for Cisco, argues that enterprise skills need packaging, evaluation, syncing, and distribution rather than scattered markdown files. He describes using skills to make cheaper models viable, converting curated TAC knowledge-base articles into maintained skills, and rolling out an eval framework across teams. He also describes syncing a repository README to Confluence with a deterministic script.
AIOpenAI is rolling out GPT-6.1 Sol Ultrafast, which generates tokens up to 8x faster than Sol Standard. On the API it costs $12 per 1M input tokens and $60 per 1M output tokens, and the main post says it is just 1.2x the cost of Astra.
AILiveKit is offering free access to its Simulations product through October, letting teams check what their agent can do and find gaps before deployment. The product also lets teams test any model against their own scenarios before switching models.
AIOpenAI is rolling out GPT-6.1 Sol Ultrafast on ChatGPT Work, Codex, and the API. The Ultrafast mode is priced at $12 per million input tokens and $60 per million output tokens, and it runs 8x faster than Sol Standard.
AIOpenAI is rolling out GPT-6.1 Sol Ultrafast, which generates tokens up to 8x faster than Sol Standard. On the API it is priced at $12 per 1M input tokens and $60 per 1M output tokens. The mode is also available today in Codex and ChatGPT Work.
AIArtificial Analysis has released Harvey LAB-AA, an evaluation built on Harvey's LAB dataset and developed in collaboration with Harvey. Full results are published on the Artificial Analysis evaluations page, alongside Harvey's commentary on the benchmark and human expert preferences.
AIArtificial Analysis reports that generating more output tokens does not necessarily yield a higher score. GPT-6 Astra (max) scored 8.6% using about 81k output tokens per task, while Grok 4.7 (xhigh) used roughly 180k yet scored lower. Three Claude models produced the most output tokens, about 202k to 562k per task, but scored between 2.8% and 6.4%.

AIAmong models with a Hallucination-Gated All-Pass Rate above 0%, GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max), and Grok 4.7 (xhigh) set the Pareto frontier for score versus cost per task. GPT-6 Luna (max) is the cheapest at about $0.22 per task, scoring 3.3%, while Grok 4.7 (xhigh) leads at about $9.50 per task and Muse Spark 1.3 (max) costs about $4.20. The three Claude models cost about $18 to $22 per task.

AIOnce hallucinations are accounted for, Muse Spark 1.3 (max) drops from 26.7% to 8.9%, leaving Grok 4.7 (xhigh) first on the headline metric at 9.4%. Kimi K3 (max) falls from 16.7% to 5.3%, and Claude Sonnet 5.5 (max with fallback) falls from 11.7% to 2.8%. GPT-6.1 Sol (max) declines least, from 7.5% to 6.9%.

AIArtificial Analysis reports that Kimi K3 (max) achieves a 93.0% Criterion Pass Rate but averages 2.09 material hallucinations per task. Muse Spark 1.3 (max) scores higher at 96.0% while averaging 1.68 material hallucinations per task, showing that completing criteria and avoiding hallucinations are distinct skills.

AIArtificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

AIAmazon Bedrock AgentCore payments lets AI agents pay for model inference one request at a time, using x402 with USDC on the Base network. Incarna used the service to connect its agents to BlockRun, a pay-as-you-go router serving more than 90 models from more than 15 providers. Spending limits are enforced at the infrastructure layer, outside the model.
AIVoyager has launched a desktop app on Mac that lets AI agents work inside creative tools such as After Effects, DaVinci Resolve, Blender, and Unity. The app reads project files, operates the apps, and outputs editable results, according to the post.
AIThe NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a smaller, clearer, and easier-to-verify agent harness. The competition required agents to answer natural-language questions over heterogeneous sources, including databases, CSV and JSON files, prose documents, PDFs, and briefing videos.
AIOpenAI has made Ultrafast mode for GPT-6.1 Sol available in all supported regions, including US and EU data residency. EU data residency has also been added for GPT-6.1 Sol Fast and GPT-6 Luna Fast. Access to Codex and ChatGPT Work is offered on Pro 500, eligible usage-based Enterprise, and credit-based Edu plans, with Enterprise admins required to enable it.
AIOpenAI's API pricing for GPT-6.1 Sol in Ultrafast mode is $12 per million input tokens and $60 per million output tokens. The mode is built for time-sensitive work such as debugging outages, agents navigating apps, and live experiences where speed matters.

AIOpenAI says Ultrafast is rolling out today for GPT-6.1 Sol in the API, Codex, and ChatGPT Work. The company describes it as near-Astra intelligence at up to 8x faster speeds than Sol Standard.
Why it matters: The post names the access points and a speed comparison to the Sol Standard tier, which helps developers judge whether the faster option fits their workflow.
AIDatabricks introduced Vibe Data Modeling, an open-source agent that helps teams build, validate, and evolve business-specific data models. It applies roughly 250 modeling rules while keeping data modelers and business stakeholders involved. Teams can start from 40 industry models as a baseline and iterate toward models that reflect how their business operates.

AIMicrosoft CEO Satya Nadella thanked President Trump for an honor bestowed at an event alongside prominent American innovators. He said he looks forward to continuing work together to advance technology and drive American prosperity. The quoted remarks from Trump credit Nadella with decades of transforming Microsoft.
AILauren Tan invites users of Grok Bot on Omarchy and developers building plugins for it to share feedback and feature requests. The post points to the Omarchy plugin catalog and asks what integrations could be supported. Background from DHH says SpaceXAI joined the Omacom Foundation as a Founding Corporate Patron, contributing $1,500,000 in Grok tokens for Omarchy's maintenance and development.
AIPhoton says its iMessage agents can now discover other agents, look up their phone numbers, and add them to a group chat with the user. The company says the feature opens agent access to every merchant and platform, and it is available in beta at photon.codes.
AITessl argues AI code review is slow because generated code outpaces trust, and proposes executable specs that let agents check preview environments against product intent. Its spec reviewer splits work between a planner agent that extracts requirements and parallel verifier agents that test each one against the code and base branch.
AITypeSafe AI's jev model was reported to be more calibrated than 40 other decision models tested on email classification. The main post itself only urges readers to stay calibrated, so no further figures or methodology are given.
AIActor Ben Affleck drew attention this week for explaining machine learning concepts, including convolutional neural networks, tensors, and transformers, in several recent interviews. He said he fine-tuned open video models by unfreezing weights and training only the last cinematic layer, using a dataset he built over about eight months for his startup. Affleck said he worries about students and learned helplessness more than Skynet, and predicted AI will be additive to the movie business.
AIArena, the crowdsourced AI model leaderboard that started as a UC Berkeley research project, raised a $200 million Series B at a $3.1 billion valuation, led by Lightspeed Venture Partners and Khosla Ventures. The company said it reached $100 million in annualized run-rate revenue in June, up from $30 million when it raised its $150 million Series A in January at a $1.7 billion post-money valuation.
AIOpenAI has reportedly told investors its annualized revenue is approaching $50 billion, about $20 billion below a previously reported $70 billion figure. The Financial Times reports the earlier number came from investor attempts to compare OpenAI with Anthropic, which counts cloud partners' sales differently. OpenAI's IPO has reportedly been pushed to early 2027.
AIGoogle announced at a Google Cloud event a unified Gemini agent that can plan and complete tasks from a single interface, starting with businesses. The agent has its own Workspace account, connects to systems including Google Workspace, Microsoft 365, Slack, and Jira through MCP, and writes an audit trail attributed to the agent. Google said consumers will get access later, after it addresses security, scale, and performance.
This story has a top pick“Google Cloud launches Gemini agent as single universal work agent”
AIThe Association of Historical Mathematicians (AHM) has called for a boycott of OpenAI after the company released more than 700 AI-generated proof files at once. Fields Medalist Terence Tao, who chairs the group, argues that AI solving open problems autonomously reduces seminars, collaborations, and fertile research directions, and that the field should shift its measure of progress toward explanation and community-building.
Why it matters: The article links the AHM boycott call to Tao's argument that AI-driven proof volume is changing how mathematicians measure progress and whether solutions remain useful.
AITrump disclosed more than 500 securities transactions in August, including a purchase of up to $25 million in Meta stock and up to $5 million in SpaceX senior unsecured notes. The filing, which reports trades in value ranges, shows total activity of roughly $74.3 million to $273.3 million according to a CNBC analysis. The SpaceX notes were bought two days before Trump signed a national space transportation policy, and the White House says the portfolio is independently managed.
AIGoogle Cloud's borderless Lakehouse lets Gemini query data on AWS and Azure without variable egress fees. It reads directly from Salesforce Data 360, SAP, ServiceNow, and Workday without copying data. It also federates open Apache Iceberg tables across Databricks Unity, Snowflake Horizon, and AWS Glue.
