Skip to contentSkip to stories

Updated

#Deployment/Engineering

Showing low-relevance items too. Hide low-relevance items

Sep 29

Sep 29Tue
  1. OpenAIOfficialAI score60

    OpenAI upgrades Codex Security Cloud with default access to cyber-capable models

    AIOpenAI says Codex Security Cloud is getting a major upgrade that includes access to cyber-capable models through Daybreak Blue by default. The upgraded tool scans entire GitHub repos, continuously reviews new commits, investigates and deduplicates findings, and prepares fixes for review even when the user's laptop is closed. It is available as a plugin in Codex desktop and web.

    Why it matters: The post names concrete capabilities, from repo-wide scanning to cloud-run fix preparation, which shows how the product changes a security review workflow.

    Video from @OpenAI's post
  2. ChatGPTOfficialAI score62

    ChatGPT Space adds shared pages for team collaboration with AI

    AIOpenAI's ChatGPT account announced ChatGPT Space, a workspace for creating and collaborating with a team and AI. Space introduces pages, interactive documents with charts, images, checklists, and dashboards that ChatGPT can build from work context or conversation.

    Why it matters: The post explains how pages, team editing, and ChatGPT tagging combine into a shared workspace, showing which plans and platforms get it first.

    Video from @ChatGPT's post
  3. OpenAIOfficialAI score62

    OpenAI makes GPT-6.1 Sol available to Plus, Pro, Business, Enterprise, and Edu users

    AIOpenAI says GPT-6.1 Sol is available starting today to all Plus, Pro, Business, Enterprise, and Edu users. The model is offered in ChatGPT Work and Codex, and the post links to OpenAI's introduction page.

    Why it matters: The post specifies which plan tiers gain GPT-6.1 Sol and in which products, showing where the new model reaches users directly.

  4. DatabricksOfficialAI score12

    Databricks argues AI needs a unified context layer, not smarter models

    AIDatabricks says AI's core limitation is missing business context, not intelligence, because data is scattered across dashboards, documents, tickets, and chats. The company promotes a unified context layer and its Genie Ontology approach, with a product demo in an on-demand webinar.

    Video from @databricks's post
  5. PerplexityOfficialAI score23

    Perplexity Computer adds Automations for ongoing recurring work

    AIPerplexity has introduced Automations in Perplexity Computer, a feature for handling ongoing work. The post links to a blog post with details, but its own text provides no further specifics on how the automations operate.

  6. Azure BlogOfficialAI score40

    SQL Server on Azure Local Becomes Generally Available for Connected and Disconnected Use

    AIMicrosoft has made SQL Server on Azure Local generally available for connected and disconnected deployments, letting organizations run SQL Server in their own datacenters and edge locations. Disconnected operations continue locally where external connectivity is restricted or unavailable. Eligible existing SQL Server licenses can be used, and Foundry Local on Azure Local, currently in preview, brings AI inference alongside SQL Server data.

  7. Meta NewsroomOfficialAI score34

    Meta Launches Forum, a Standalone App for Browsing Facebook Groups

    AIMeta is testing Forum, a standalone iOS and Android app in the US that syncs with users' Facebook Groups to consolidate their conversations in one place. The update adds a new top-contributor role replacing previous badges, an AI-powered Ask feature that surfaces group posts and comments, and topic labels for exploring interests.

  8. Replit BlogOfficialAI score62

    Replit Agent lets the core model choose subagents and effort instead of a router

    AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.

    Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.

  9. Alex HeathXAI score34

    Factory CEO Matan Grinberg says AGI is already here

    AIFactory CEO Matan Grinberg, whose AI coding startup builds Droid agents, argues AGI is already here and explains why the company bets on many competing models. The discussion covers balancing model performance against token costs and why companies should avoid depending on a single AI provider. It also touches on hiring, the open-versus-closed AI debate, and competition with Cognition.

    Video from @alexeheath's post
  10. PerplexityOfficialAI score60

    Perplexity open-sources Bumblebee to scan developer machines for risky packages

    AIPerplexity has open-sourced Bumblebee, a read-only scanner for macOS and Linux that checks developer machines for risky packages, extensions, and AI tool configurations. When connected to Computer, it can trigger deeper scans whenever a new supply-chain risk emerges. The post says Computer reviews findings from Bumblebee and Numbat to propose better detection rules, and humans approve every change before it ships.

    Why it matters: The post shows how a read-only scanner fits into a human-approved pipeline that updates detection rules after supply-chain risks emerge, useful for teams planning developer machine security.

  11. DatabricksOfficialAI score22

    Databricks rolls out frontier models to employees on Day 1 via Unity Gateway

    AIDatabricks says it aims to give its employees the best models on launch day, quickly adopting new releases such as Opus 5.5 and GPT-6 Sol while tracking real-world usage and cost. Its AI engineering team uses Unity Gateway to manage access, spend, and model selection across thousands of employees, and to decide which models join its AI stack.

    Image from @databricks's post
  12. Microsoft ResearchOfficialAI score75

    Microsoft Research introduces Quine, a multimodal biology world model and research harness

    AIMicrosoft Research introduced Quine, an experimental research system combining a multimodal world model of biology with an interactive harness that connects models, scientific tools, literature, and researchers. In a pancreatic cancer study with the Broad Institute, Quine prioritized compounds that shifted tumor cell states, and several top-ranked candidates were validated in wet-lab assays. Access is initially limited to the Quine Fellows program and select collaborations, and the system is intended for research use only, not clinical use.

    Why it matters: The post shows how a multimodal biology world model is wired into a harness, grounded in one wet-lab cancer example and a limited fellows-program access path.

  13. Ahead of AI (Sebastian Raschka)BlogAI score43

    Language Models for Text Classification: From Bag-of-Words to Jev

    AISebastian Raschka traces text classification from bag-of-words models such as naive Bayes and logistic regression through pre-transformer neural networks, then sets up an analysis of the recently released Jev AI model. The article frames Jev as a general-purpose classifier that trades specialized accuracy for speed, cost, and breadth of tasks.

  14. IEEE Spectrum · AINewsAI score14

    IC-STAR Brings Full-Flow Autonomous AI to Digital and Analog Chip Design

    AIThe webinar presents IC-STAR, an autonomous AI approach that shifts silicon engineers from manually managing tools and handoffs to defining objectives and supervising AI-driven execution across the chip development lifecycle. It covers four enabling technologies and includes a look at Ambiq's production deployment of autonomous AI. The source provides no performance figures or availability details.

  15. ModelScopeOfficialAI score44

    Intern-Decision multimodal models scale structured decisions at 0.8B–4B

    AIShanghai AI Laboratory's Intern-Decision family of 0.8B, 2B, and 4B multimodal models averages 79.38, 84.68, and 90.02 across seven decision benchmarks. Intern-Decision-4B scores 88.74, surpassing Jev while achieving better probability calibration. Reported mean latency is 33.98, 33.28, and 44.16 ms, versus 109.70 ms for Jev in the same local HF setup.

    Image from @ModelScope2022's post
  16. OpenBMBOfficialAI score34

    MiniCPM-o 4.5 now runs in SGLang Omni v0.1.7 for developers

    AIOpenBMB announced that MiniCPM-o 4.5 is now supported in SGLang Omni v0.1.7, giving developers more flexibility to run and build with the model. The background release notes add that MiniCPM-o 4.5 brings multimodal input and speech output to the runtime. MiniCPM-o and MiniMax-Music3 also gained Intel XPU support in the same release.

  17. vLLMOfficialAI score58

    IQuest-Q1 320B MoE coding model gets day-0 support in vLLM

    AIvLLM announced day-0 support for IQuest-Q1, a 320B-parameter MoE model with 15B active per token, 256 experts with 8 active, and a 524,288-token context. The post credits existing vLLM features such as the hybrid KV cache coordinator, sinks attention path, and EAGLE speculative decoding with probabilistic draft sampling. The linked material includes a Docker image and vllm serve commands, with and without recursive MTP.

    Image from @vllm_project's post
  18. Azure BlogOfficialAI score75

    Microsoft announces Fabric IQ in Copilot, Power BI agentic app creation, and new Fabric and SQL updates

    AIMicrosoft announces new Microsoft Fabric and SQL Server updates at FabCon and SQLCon in Barcelona, including Fabric IQ integration with Microsoft Copilot Chat and Cowork, now generally available. Power BI agentic app creation enters preview in the coming weeks for Pro and Premium Per User customers, with Fabric Apps database capabilities up to 1 GB per app at no additional cost.

    Why it matters: The post lists dozens of Fabric and SQL updates tied to Copilot and agents, with specific availability and pricing terms for Power BI customers that help readers judge what applies to them.

  19. Thomas WolfXAI score29

    Thomas Wolf calls a post simply "impressive"

    AIThomas Wolf, owner of the Hugging Face account, posted the single word "impressive" in response to a quoted post. The quoted post reports a new NanoGPT training record of 39.9s, down 27.7s from the prior 67.6s, achieved through per-flop optimizations such as sampled softmax and sparse updates.

  20. SGLangOfficialAI score53

    SGLang adds Day-0 support for IQuest-Q1 with a single-node serve command

    AISGLang says it has Day-0 support for IQuest-Q1, an open-source sparse MoE model with 320B total and 15B active parameters for coding and agentic tasks. The post includes a single-node serving command for H200 GPUs in BF16, using tensor parallelism of 8, EAGLE speculative decoding, and the iquest_q1 reasoning and tool-call parsers. The image marks the command as not verified.

    Image from @sgl_project's post
  21. Matei ZahariaXAI score36

    Matei Zaharia says autoresearch results are going into serving stack

    AIMatei Zaharia said autoresearch produced strong results that are being integrated into a model serving stack. The post gives no specific figures, benchmarks, or product names. Background context from a related post says Databricks ranked #1 on NVIDIA's SOL-ExecBench kernel leaderboard across all four tracks using agents.

  22. howie.seriousXAI score13

    Claude Code account bans may stem from poor IP quality

    AIThe post says Claude Code account bans often result from low-quality IP addresses, and recommends the IPCheck.ing tool to check IP quality. It adds that many promoted home broadband and VPS services should be verified with such a tool before trusting them. Background from @jason5ng32 says IPCheck.ing updated its IP quality scoring to v3 and that one VPS provider widely recommended on X scored 54 in a sampled IP check.

  23. howie.seriousXAI score23

    Qwen3-VL 32B tags 40,000 Eagle images on a 128GB Mac

    AIA user ran qwen3-vl:32b-instruct locally through Ollama on a 128GB computer to auto-tag 40,000 images in their Eagle app. The image library was collected over 10 years and is intended as a personal asset base for Claude Code-made knowledge videos. The user says the use case makes the 128GB memory purchase feel worthwhile.

    Image from @howie_serious's post
  24. Mastra BlogOfficialAI score42

    Mastra Adds Memory Hooks to Observe and Modify Agent Memory Cycles

    AIMastra has added memory hooks that let developers monitor or alter an agent's observational memory cycles. Lifecycle hooks such as onObservationStart and onReflectionEnd report on each cycle, including token usage for spotting cost spikes, while transform hooks like beforeObservation and afterReflection can prune, remove, or redact memory data.

  25. Artificial Analysis ArticlesOfficialAI score62

    Artificial Analysis open-sources AA-AgentPerf-Local for benchmarking local AI agents

    AIArtificial Analysis has open-sourced AA-AgentPerf-Local, a tool that replays recorded agent trajectories to measure inference speed on laptops and workstations. Initial results cover NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro, with the RTX 5090 fastest for models that fit its 32 GB. The source states the tool and leaderboard will expand to more hardware, frameworks, and models.

    Why it matters: The source gives per-system completion times and memory bandwidth figures, letting readers compare local hardware for running agentic workloads.

  26. Manus BlogOfficialAI score50

    Manus Flex lets users connect their own API keys to the Manus workspace

    AIManus is launching Manus Flex, a module that lets users power Manus agents with their own API key from a supported inference provider. Model inference is billed directly by that provider, while other services used in Manus tasks still consume Manus credits. OpenRouter, Fireworks, and Modal are announced as initial inference partners for the Flex Inference Partner Program.

  27. Suno BlogOfficialAI score12

    Three Essential Tips for Using EQ in Music Production

    AIEqualization (EQ) is one of the most widely used music production tools, and this guide offers three tips for using it well. The advice covers mixing by ear rather than by the visual curve, cutting problem frequencies before boosting, and placing EQ first in the effects chain so later effects process a cleaner signal. Suno Studio's per-track EQ supports multiple EQs per track and sharing of presets.

  28. Luma AI NewsOfficialAI score22

    AI Photo Editing Prompt Formula Preserves Color, Light, and Skin in Campaign Edits

    AIThe article presents a four-part prompt structure (action verb, target element, desired result, protection instructions) for AI photo editing, saying it preserves approved work across platforms. It identifies three common failure causes: unmatched light direction, stacked edits in one prompt, and vague visual language. It states that simple skin retouching takes 2-3 minutes versus 15-30 minutes manually.

Sep 28

Sep 28Mon
  1. ModelScopeOfficialAI score44

    Audio8 ASR Infinite enables unlimited-length streaming speech transcription with bounded memory

    AIAudio8 ASR Infinite transcribes Chinese and English audio of unlimited length using a rolling KV Cache that avoids accumulated drift. At a 480 ms delay, it reports 1.75 CER on AISHELL-1, 2.89 on AISHELL-4, and 3.04/6.81 WER on LibriSpeech test-clean/test-other. The preview release is under Apache 2.0, with deployment through an adapted vLLM stack.

    Video from @ModelScope2022's post
  2. clem 🤗XAI score23

    AMD acquires World Labs, led by Fei-Fei Li, for AI world models

    AIAMD is bringing World Labs and Fei-Fei Li into its organization, according to Lisa Su's post, which says the combination will pair World Labs' AI and world-model expertise with AMD's compute leadership. Hugging Face co-founder Clément Delangue congratulated the team and said he looks forward to what they will build in the coming years.

  3. KrASIA · Big TechNewsAI score47

    Alibaba unveils Zhenwu V900 AI chip, targets 20 GW data center capacity by 2032

    AIAlibaba unveiled the Zhenwu V900 AI chip at its 2026 Apsara Conference, claiming three times its predecessor's performance and support for clusters of up to 500,000 cards. The company is pursuing data center capacity beyond 20 gigawatts by 2032 and has committed RMB 380 billion in capital spending over three years.

  4. KreaOfficialAI score22

    Seedance 2.5 Draft Mode now available on Krea

    AIKrea has launched Draft Mode for Seedance 2.5, letting users experiment with 480p generations before switching to 1080p once a scene is right. The post directs readers to try the feature on Krea's platform.

    Video from @krea_ai's post
  5. vLLM BlogOfficialAI score54

    vLLM guide explains disaggregated serving for prefill and decode

    AIThe vLLM blog guide explains how separating prefill and decode, and moving tokenization to a CPU-only render tier, can keep token streams from stalling under load. In a two-L40S test on Qwen2.5-7B, collocated p99 inter-token latency reached 169 ms at 0.4 req/s while disaggregated serving stayed between 25 and 52 ms. The guide notes that the gain depends on fast KV cache transfer, and it includes setup code for NIXL-based serving and the render/derender API.

  6. Amp NewsOfficialAI score34

    Amp Adds Plaid Speed for GPT-6 Astra Modes at 6x Speed and Cost

    AIAmp now supports Plaid speed for modes that use GPT-6 Astra, using OpenAI's ultrafast tier to run inference up to 6× faster at 6× cost per token. Plaid works only with Amp-provided inference, not linked ChatGPT subscriptions, and subagents and non-Plaid inference fall back to fast or standard speed.