Explore Community Presets for Higgsfield Katana.
AIMotion graphics, 3D animations, product launches, fashion, car, travel, and aura-farming edits. Pick a preset, add your characters, products, or clothes, and recreate it in Claude. More coming soon.
Updated
Updated
AIMotion graphics, 3D animations, product launches, fashion, car, travel, and aura-farming edits. Pick a preset, add your characters, products, or clothes, and recreate it in Claude. More coming soon.
AIHong Kong's Centre for Artificial Intelligence and Robotics (CAIR), Chinese Academy of Sciences, unveiled CARES 4.0, a multimodal clinical AI agent that carries out tasks rather than only answering questions about images or video. Built on the Harness agent framework and CAIR's own CT, MRI, ultrasound, endoscopy and EEG foundation models, it has been validated at several top-tier hospitals. CAIR says the system gives suggestions with reasoning paths and sources, and that the doctor remains the final decision-maker.
AIMirroS released AgentGarten, which pairs executable code environments with a real-time neural renderer running above 30 fps so agents can act, observe, and learn. In a one-on-one hide-and-seek setup, the hider learned to block passages by round 4 and the seeker learned to climb ramps by round 10, guided by notes the agents wrote after each round. The authors report applying the same loop to four other tasks, including a dog-companion game, a narrow-bridge car passing task, herding, and quarry loading.
AIGoogle Cloud introduced the Gemini agent, a general office agent that can search, write emails, build slides, analyze data, run code, and coordinate sub-agents. It can take on an enterprise identity with email, calendar, and account, and it selects underlying models automatically, including Anthropic's Claude. The article presents this alongside OpenAI's Dots and Meta's Muse as competing office and personal agents.
AIShanghai AI Lab has open-sourced Intern-Decision, a family of 0.8B, 2B and 4B parameter models that return structured decisions with probabilities instead of free text. The developers self-report that the 4B model averages 90.02% accuracy across seven test suites, ahead of a commercial reference model at 88.74%, with about 44 milliseconds of local latency on a single RTX 4090. Weights are on Hugging Face, and MetaX says the models run on its hardware from launch.
AIAt the Apsara Conference, Alibaba's Qwen team outlined a roadmap of Qwen4 followed by Qwen4.5 and Qwen5, aiming for 5T to 10T parameters. The article notes that Qwen3.8 reached 2.4T parameters and that Qwen3.8-Flash activates 6B parameters per inference while cutting training cost to one-ninth. It also describes Qwen3.8-Max running model-driven experiments in chip design and inference optimization, and multimodal updates including a video model slated for November.
AIOpenAI is rolling out GPT-6 in ChatGPT with an "Intelligent UI" that presents answers as interactive graphics, buttons, charts, and forms instead of plain text. Users can also build small tools such as a savings calculator or a game inside the chat. The rollout starts globally today for Plus, Pro, Business, and Enterprise users, with Free and Go users following one day later.
AIOpenAI is rolling out Intelligent UI in ChatGPT, which integrates interactive visuals such as tappable buttons, charts, editable graphs, and custom calculators into conversations. The feature launches with GPT-6 for Pro, Plus, Business, and Enterprise users, arrives Thursday for free and Go tiers, and lets users reduce how many visuals appear.
AIOpenAI is rolling out an Intelligent UI feature in ChatGPT that lets answers combine text with diagrams, charts, forms, and tappable buttons. It is available starting today to Plus, Pro, Business, and Enterprise users, and will expand to Go and free tiers on Thursday. Higher-tier users get the mid-range GPT-6 Sol model, while Go and free users get GPT-6 Luna.
AIXiaomi released the MiMo-V2.5-TTS Series, three speech synthesis models for stock voices, voice design, and voice cloning. The models accept natural-language style instructions and inline audio tags, and the source says the three models are free of charge for a limited time on the Xiaomi MiMo API platform. Xiaomi also open-sourced integration Skills for agent applications on GitHub.
Why it matters: The release shows how a TTS family adds style instructions, inline audio tags, and voice design or cloning to speech synthesis, which matters for agent and creative workflows.
AIGoogle published a prospective study of AMIE, a research conversational system that patients chat with before doctor appointments, in The Lancet with Beth Israel Deaconess Medical Center. Clinicians reported the summaries helped them prepare for visits in 75% of cases and influenced their approach to care in more than half. AMIE's differential diagnoses matched the doctors' final diagnoses 90% of the time.
Why it matters: The study tests a patient-facing diagnostic chat system in a real urgent care clinic, a setting that goes beyond lab evaluation and is useful for judging clinical readiness.
AIArtificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier on AA-Video-T2V v2.0 in Multi-Scene & Narrative, Lighting & Materials, and Text Rendering. It is furthest behind in Dialogue & Lip Sync and Human Anatomy. Compared with Grok Imagine Video 1.5, Lite matches it in Physics and trails on the other nine capabilities, by the least in Multi-Scene & Narrative.
AIDevelopers are pairing frontier AI models with NVIDIA Omniverse libraries to turn simulation ideas into working applications, from humanoid warehouse simulators to autonomous-driving test environments. In the examples, developers direct AI agents through natural-language instructions and review results, while Omniverse provides GPU-accelerated physics, rendering and sensor simulation. One experiment reported a simulated Unitree G1 humanoid clearing a hurdle in 64 of 100 trials.
AIHospitals, academic centers, medtech firms, and pharma companies all face the same obstacle: imaging data is locked in clinical systems and hard to share. The EXAM study across 20 institutions showed federated learning, which shares model weights rather than patient data, improved AUC by 16% on average. Collaboration remains difficult due to scanner and protocol heterogeneity, privacy governance, and the lack of a common data substrate.
AIOur most powerful AI video editing tool, inside Claude. Upload a reference and create editable motion graphics, product launch videos, or aura-farming edits. Available now in Claude via Higgsfield MCP.
AIUse Claude Motion to animate charts, create customer walkthroughs or create short explainers, then bring your work into Runway to generate videos and images. Learn more at the link below.
AIAnthropic launched two beta features for Claude: Dashboards, which turns connected data sources like BigQuery, Databricks, Snowflake, or Salesforce into auto-updating live dashboards from text prompts, and Motion, which creates animated explainer videos from text, diagrams, and images. Dashboards is available to paid users and Motion to Team and Enterprise plans, while Docs, Slides, and Design leave beta and work across all plans, including free accounts.
AIClaude Docs, Slides, and Design are out of beta and available on every plan, including Free, starting today. The source also says a team and Claude can edit the same doc, deck, or design together.
AIAsk Claude to turn your data into live dashboards and your ideas into animated explainers.
AIConnect with Sahil Dua and explore an open multimodal model that unifies text, images, audio, and video representations. #COLM2026 @GoogleDeepMind
AI…robots, power humanoids, drive cars, fly drones, and other autonomous machines.
AIWe're particularly excited by agents learning from experience inside Odyssey-3, working to accomplish objectives.
AIIt's a new kind of world simulator, and learns representations of physics, dynamics, and cause-and-effect from a broad dataset of visual observations.
AI…ComfyUI integrations plus local model support on up to 128 GB of unified memory. A complete creative stack that runs on your own machine.
AITheir new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it:
AIsuper simple: llama serve -hf ggml-org/Clef-Flash-GGUF browse all the decision models here we also polished Llama App website & docs 🌟
AIMeta is donating 1,000 Ray-Ban Meta AI glasses to four Singapore organisations serving people with disabilities, alongside a US$30,000 grant for accessibility training. The glasses help users who are blind or have low vision read text, identify objects and describe their surroundings. The grant will fund a free curriculum from the Singapore Association of the Visually Handicapped on using the glasses safely in daily life.
AIPerplexity has released pplx-embed-v2-late, a pair of ColBERT-style multimodal embedding models in 0.6B and 9B sizes that retrieve text, images and rendered PDF pages in a shared embedding space. Both are available on Hugging Face under the MIT license, while a hosted API endpoint is planned but not yet live.
AIJohns Hopkins astrophysicist Brice Ménard, working as an Anthropic researcher, used Claude Science to produce the first complete map of the sky in ultraviolet light. Claude orchestrated agents to merge GALEX, Swift, and FIMS/SPEAR data, then predicted roughly a third of the sky that no UV telescope had observed, using relationships to visible, infrared, and radio data. Hidden test regions were reconstructed to within about 10% of real measurements, and each pixel is labeled measured or predicted with uncertainty estimates.
Why it matters: The post shows how an astrophysicist used Claude Science agents to merge UV surveys and predict missing sky regions, with a validation step that makes the method reusable.
AILuma announced that Claude Motion animations can now open directly in Luma through an MCP connection, letting creators restyle them and reframe them to 9:16, 1:1, 4:3, or 21:9. Claude Motion, currently in beta on Claude Team and Enterprise plans, generates animated explainers from prompts, while Luma's Ray and Uni video models produce final video files.
AIA NVIDIA study accepted at NeurIPS 2026 reports that multimodal models refuse harmful requests less reliably when they call tools. Refusal failures rise by up to 68.7% relative and by 17.7% on average across the models tested, including Claude Opus 4.6 and 4.7 and Gemini Agentic Vision. The authors attribute this to tool outputs crowding out the original harmful intent and to attention shifting toward describing tool results. Re-inserting the original request and image before the final response restores part of the lost refusals.
AIThe vLLM team released a technical report on vLLM-Omni, a unified serving runtime for omni-modality generation spanning multi-stage autoregressive pipelines, iterative diffusion, and stateful sessions. Current LLM servers and diffusion stacks each cover only one of these patterns, pushing deployments to stitch disjoint runtimes together. vLLM-Omni offers a shared control plane in which an orchestrator advances requests across stages, specialized engines handle compute, and a connector carries payloads.
AIThis Next Token episode discusses Personal Agents, including Dots in Codex, memory and cloud computer permissions, and whether agents should act as assistants or digital twins. The hosts also cover Instinct's booking and business-travel model, hands-on projects built with Opus 5.5, and whether software, games, and hardware could become open source as AI makes rewriting easier.
AIApple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as a conditional normalizing flow trained with exact likelihood. The model matches or outperforms strong image generation baselines on text-to-image benchmarks in just four sampling steps while retaining exact likelihood over the generative trajectory.