Skip to content

#Image generation

Oct 8

TodayOct 8Thu14 items
  1. Artificial AnalysisAI score38

    Grok Imagine Video 1.5 Lite leads in architecture, consumer, and knowledge-work use cases

    Artificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier in Architecture & Real Estate, Consumer, and Productivity & Knowledge Work use cases. It sits furthest from the frontier in Live-Action Film and Frontier use cases. Against Grok Imagine Video 1.5, Lite matches it in Social Media & Creator Content and trails it on the other nine use cases.

  2. Artificial AnalysisAI score31

    Grok Imagine Video 1.5 Lite sits on the quality and speed frontier on AA-Video-T2V-Silent v2.0 Among the 12 models on AA-Video-T2V-Silent v2.0 that we benchmark for generation speed, no model is both faster and higher quality than Grok Imagine Video 1.5 Lite. It generates a 10 second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher and takes 94 seconds for a 5 second clip. Vidu Q3 Turbo is 9 seconds faster on a 5 second 720p clip, and scores well below it.

    Grok Imagine Video 1.5 Lite sits on the quality and speed frontier on AA-Video-T2V-Silent v2.0 Among the 12 models on AA-Video-T2V-Silent v2.0 that we benchmark for generation speed, no model is both faster and higher quality than Grok Imagine Video 1.5 Lite. It generates a 10 second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher and takes 94 seconds for a 5 second clip. Vidu Q3 Turbo is 9 seconds faster on a 5 second 720p clip, and scores well below it.

  3. MidjourneyAI score34

    We're testing a new "thinking mode" for our image generation on our Alpha website (alpha dot midjourney dot com). We're finding it boosts prompt accuracy, typography and coherence. We'd love your help testing it on your images and telling us what you think. Thanks! <3

    We're testing a new "thinking mode" for our image generation on our Alpha website (alpha dot midjourney dot com). We're finding it boosts prompt accuracy, typography and coherence. We'd love your help testing it on your images and telling us what you think. Thanks! <3

  4. Midjourney UpdatesAI score46

    Midjourney Tests Thinking Mode for Image Generation on Alpha Site

    Midjourney is testing a "Thinking Mode" on its Alpha website, where users can click "Rerun (Thinking)" in a job's lightbox to regenerate an image. The company says early tests show gains in prompt accuracy, typography, and coherence, and it is asking users to share feedback in its #ideas-and-features channel. It may later offer the mode broadly or as an option to add more thinking after a job.

  5. Midjourney UpdatesAI score25

    Midjourney Adds Shared Folders and Thinking Mode in Alpha Update

    Midjourney's alpha site now lets users share folders with others as Collaborators or Viewers, with sharing by link also available. A new thinking mode lets users rerun jobs made with 8.2 standard and edit models to fix missed prompt details such as objects, layout, anatomy, and text. Sharing does not change image privacy, so non-stealth images can still appear on Explore and profiles.

  6. RunwayAI score36

    Bring your ideas to life with @claudeai Motion + Runway. Use Claude Motion to animate charts, create customer walkthroughs or create short explainers, then bring your work into Runway to generate videos and images. Learn more at the link below.

    Bring your ideas to life with @claudeai Motion + Runway. Use Claude Motion to animate charts, create customer walkthroughs or create short explainers, then bring your work into Runway to generate videos and images. Learn more at the link below.

  7. ClaudeAI score46

    Claude Motion turns reports, charts, or product walkthroughs into short animations. Claude writes each one as code, with no video model involved, so you can change any word, number, or timing, then export an MP4. In beta on Team and Enterprise plans.

    Claude Motion turns reports, charts, or product walkthroughs into short animations. Claude writes each one as code, with no video model involved, so you can change any word, number, or timing, then export an MP4. In beta on Team and Enterprise plans.

  8. Elvis SaraviaAI score42

    Don't sleep on domain-specific harnesses. Coding agents are great because of their harnesses, but they aren't built for creative work. Creative work needs its own harness. Voyager looks great. It's an open harness for video, graphics, and games. The agent works with your files and drives apps like Blender, DaVinci Resolve, and Unity right on your desktop. Bring Opus, Astra, or DeepSeek. Excited to try this one.

    Don't sleep on domain-specific harnesses. Coding agents are great because of their harnesses, but they aren't built for creative work. Creative work needs its own harness. Voyager looks great. It's an open harness for video, graphics, and games. The agent works with your files and drives apps like Blender, DaVinci Resolve, and Unity right on your desktop. Bring Opus, Astra, or DeepSeek. Excited to try this one.

  9. QbitAI (้‡ๅญไฝ)AI score47

    Vidu Q4 Preview Offers 4K Video Generation at About 0.09 Yuan per Second

    Shengshu Technology has opened a preview of its Vidu Q4 video generation model, which supports native 4K output and up to 15 reference images and three reference audio clips. Testers generated a one-minute video for about 5.4 yuan, roughly 0.09 yuan per second at 720P, which the article says is a starting price that varies by resolution and mode. The Vidu Q4 preview is available through the Vidu platform, with the MaaS API priced at about 0.6 yuan per second for 720P image-to-video.

Oct 7

Oct 7Wed
  1. Apple Machine Learning ResearchAI score42

    Apple's Normalizing Trajectory Models generate images in four steps with exact likelihood

    Apple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as a conditional normalizing flow trained with exact likelihood. The model matches or outperforms strong image generation baselines on text-to-image benchmarks in just four sampling steps while retaining exact likelihood over the generative trajectory.

  2. Amazon ScienceAI score35

    AI-powered smart glasses that guide delivery drivers from van to the customer's doorstep, hands-free. Hear from Amazon applied scientist @_yelinkim on the computer vision and edge AI behind it at @COLM_conf. #COLM2026

    AI-powered smart glasses that guide delivery drivers from van to the customer's doorstep, hands-free. Hear from Amazon applied scientist @_yelinkim on the computer vision and edge AI behind it at @COLM_conf. #COLM2026

  3. Elvis SaraviaAI score22

    AI image tools have gotten very good at giving you exactly what you ask for. Knowing what to ask for is still hard. @envato just launched Burst Mode to help. You start from a rough idea or a reference image, get up to 6 directions side by side, then steer them with moodboards and a creativity control that goes from focused to wild. A really good experience and a great way to get strong results fast.

    AI image tools have gotten very good at giving you exactly what you ask for. Knowing what to ask for is still hard. @envato just launched Burst Mode to help. You start from a rough idea or a reference image, get up to 6 directions side by side, then steer them with moodboards and a creativity control that goes from focused to wild. A really good experience and a great way to get strong results fast.

  4. falAI score34

    Vidu Q4 is now live on fal - Image-to-video from a single first frame, with native audio - Reference-to-video with up to 12 reference images and 3 voice clips for consistent characters and voices - 3 to 16 second clips from 540p up to 4K

    Vidu Q4 is now live on fal - Image-to-video from a single first frame, with native audio - Reference-to-video with up to 12 reference images and 3 voice clips for consistent characters and voices - 3 to 16 second clips from 540p up to 4K

  5. SantiagoAI score32

    Really cool feature for anyone who wants to design something but isn't sure what to do: 1. Start with an idea 2. Use burst mode 3. Get 6 images back 4. Extremely fast From there, you can explore a specific direction by clicking on "more like this", or zero in on a specific style by clicking on "polish".

    Really cool feature for anyone who wants to design something but isn't sure what to do: 1. Start with an idea 2. Use burst mode 3. Get 6 images back 4. Extremely fast From there, you can explore a specific direction by clicking on "more like this", or zero in on a specific style by clicking on "polish".

  6. Google DeepMind ยท The KeywordAI score62

    Google expands SynthID Detector globally to check AI-generated media

    Google is making its SynthID Detector available globally in English, letting anyone check whether an image, video, or audio file was made with AI from Google or partners including OpenAI, NVIDIA, Kakao, and soon Apple. The tool joins built-in verification in Search, the Gemini app, and Chrome, which now handle over 1 million requests daily. Google says SynthID has watermarked over 180 billion images and videos and 240,000 years of audio.

    AIWhy it matters: The source specifies which vendors' AI media the detector checks, helping readers judge how far the verification covers content they encounter online.

  7. PixVerseAI score20

    Your next video starts in the chat. The PixVerse plugin lets you create from text, images, or video references right where youโ€™re working on the idea. Hereโ€™s how to try it: โ†’ Select the PixVerse plugin. โ†’ Describe a scene, add an image, or provide a video reference. โ†’ Specify the model, duration, resolution, and aspect ratio. โ†’ Ask it to generate.

    Your next video starts in the chat. The PixVerse plugin lets you create from text, images, or video references right where youโ€™re working on the idea. Hereโ€™s how to try it: โ†’ Select the PixVerse plugin. โ†’ Describe a scene, add an image, or provide a video reference. โ†’ Specify the model, duration, resolution, and aspect ratio. โ†’ Ask it to generate.

  8. SenseTimeAI score13

    SenseNova 6.8 Flash Preview powers Dynamic Design animation

    A new angle on creativity. ๐——๐˜†๐—ป๐—ฎ๐—บ๐—ถ๐—ฐ ๐——๐—ฒ๐˜€๐—ถ๐—ด๐—ป, ๐—ฝ๐—ผ๐˜„๐—ฒ๐—ฟ๐—ฒ๐—ฑ ๐—ฏ๐˜† ๐—ฆ๐—ฒ๐—ป๐˜€๐—ฒ๐—ก๐—ผ๐˜ƒ๐—ฎ ๐Ÿฒ.๐Ÿด ๐—™๐—น๐—ฎ๐˜€๐—ต ๐—ฃ๐—ฟ๐—ฒ๐˜ƒ๐—ถ๐—ฒ๐˜„, ๐—ฏ๐—ฟ๐—ถ๐—ป๐—ด๐˜€ ๐˜€๐˜๐—ฎ๐˜๐—ถ๐—ฐ ๐—ถ๐—บ๐—ฎ๐—ด๐—ฒ๐˜€ ๐˜๐—ผ ๐—น๐—ถ๐—ณ๐—ฒ.

Oct 6

Oct 6Tue
  1. meng shaoAI score52

    xAI Cookbook adds five apps, expanding Grok API examples to ten

    The xAI Cookbook now has ten runnable Grok API examples across three tracks: real-time voice agents, multimodal generation, and live X data analysis. The author says four voice examples show the same Realtime Voice API across WebSocket, WebRTC, Twilio phone, and mobile transports. The four multimodal examples chain understanding, image generation or editing, video, and TTS, with Grok making creative decisions and Imagine models executing them.

  2. Comfy BlogAI score43

    Gemini Nano Banana 2.1 Is Now Available via ComfyUI Partner Nodes

    Google's Gemini Nano Banana 2.1 image generation and editing model is now available through ComfyUI Partner Nodes, succeeding Nano Banana 2 with a balance of price and performance. It accepts up to 14 reference images, outputs images up to 4K, and offers Minimal, Medium, and High thinking levels plus search grounding and 9:21 aspect ratio support.

  3. falAI score38

    Generate images and videos directly in Claude or ChatGPT Codex using the new fal Claude Connector and GPT plugin. No API Key required! Watch the tutorial on YouTube: https://youtu.be/gSFCIxZSodQ?si=Ti7umkZA7tClphMD

    Generate images and videos directly in Claude or ChatGPT Codex using the new fal Claude Connector and GPT plugin. No API Key required! Watch the tutorial on YouTube: https://youtu.be/gSFCIxZSodQ?si=Ti7umkZA7tClphMD

  4. falAI score34

    fal is now available in ChatGPT and Codex. Generate media with fal, then view and use it without leaving your chat. - Browse generated images and videos directly in the conversation - Access your fal Media Library from ChatGPT - Pin your media library to the sidebar to quickly bring assets into any chat Your creative assets, now where you work.

    fal is now available in ChatGPT and Codex. Generate media with fal, then view and use it without leaving your chat. - Browse generated images and videos directly in the conversation - Access your fal Media Library from ChatGPT - Pin your media library to the sidebar to quickly bring assets into any chat Your creative assets, now where you work.