
Grok Imagine Video 1.5 Turns xAI Into a Real Video Workflow Player
xAI moved Grok Imagine Video 1.5 out of preview, launched a faster app model, and added workflow features around Projects, multiple agents, and search. The release does not dethrone Seedance on Artificial Analysis, but it makes Grok a credible image-to-video product surface rather than just a benchmark entry.
The change: Grok Imagine 1.5 moves from demo to product surface
xAI has taken Grok Imagine Video 1.5 out of preview and put it into the Imagine API as
grok-imagine-video-1.5; the company also rolled out Video 1.5 Fast on grok.com/imagine plus the iOS and Android Grok apps. The official post is dated June 16, 2026, and xAI amplified it on X on June 17 with the claim that the image-to-video model brings "sharper realism, better physics and faster generations" 12.Loading content card…
This is a different kind of release from a pure leaderboard jump. Grok Imagine 1.5 Preview was already visible in arena rankings. What changed is distribution: xAI now has a generally available API model, a faster consumer app model, and a set of workflow features meant to make Grok Imagine less like a one-off clip generator and more like a lightweight production workspace 1.
What actually changed
The short version: 1.5 is an image-to-video upgrade focused on audio, motion stability, and latency. xAI says sound effects, ambience, and dialogue are generated in the same pass and better synchronized with action; it also claims movement holds together longer, with fewer warps and more believable weight and momentum 1.
| Change | xAI's stated update | Why it matters for users |
|---|---|---|
| API availability | grok-imagine-video-1.5 is out of preview in the xAI API, with image input, motion prompting, resolution, and duration controls 1 | Developers can now build around a named model rather than a preview SKU that may disappear. |
| Speed | Video 1.5 Fast generates 6-second, 720p videos in about 25 seconds, down from more than 40 seconds on the previous model 1 | Latency is a workflow variable. A 25-second retry loop changes how aggressively teams can iterate on shots. |
| Audio | xAI says effects, ambience, and dialogue are generated in the same pass and land on the action 1 | Native audio matters because silent video models force a second editing step and often break timing. |
| Workspace features | Projects, multiple agents, and library search are being rolled out over the next few days 1 | xAI is trying to own the creative workspace, not just the generation endpoint. |
The API landing page also frames Imagine as a broader visual-generation API: image and video generation, editing, restyling, product placement, virtual try-on, and mockups-to-reality. It advertises video clips up to 15 seconds and images up to 2K resolution, with image pricing starting at $0.02 per image 3. That breadth matters because video models are increasingly sold as part of a visual asset pipeline, not as isolated prompt boxes.

Where it sits in the race
The current Artificial Analysis image-to-video leaderboard with audio still has ByteDance's Dreamina Seedance 2.0 720p in first place at 1,194 Elo. Grok Imagine Video 1.5 Preview is second at 1,112 Elo, ahead of HappyHorse 1.0 at 1,091, Veo 3.1 at 1,089, and several Kling 3.0 variants 4.

That ranking cuts both ways. It gives xAI credible third-party validation: Grok's preview model is not a toy, and it is competitive in image-to-video with audio. But it also shows the gap to Seedance remains large in that category, roughly 82 Elo points on the snapshot I fetched 4.
There is also a category distinction that matters. Artificial Analysis lists Grok Imagine Video 1.5 Preview on the image-to-video leaderboard, while the text-to-video with audio leaderboard still lists the older
grok-imagine-video at 1,069 Elo, behind Seedance, HappyHorse, SkyReels, Kling, Veo, Vidu, and PixVerse entries 5. The 1.5 announcement is therefore strongest when read as an image-to-video and workflow release, not as proof that xAI has taken the overall text-to-video crown.Why speed may matter as much as rank
For creators and product teams, video generation cost is only half the pain. The other half is waiting through failed attempts. A model that returns a 6-second 720p clip in about 25 seconds changes the number of variants a user is willing to try before abandoning the shot 1.
That is especially relevant for image-to-video. Many teams already have product photography, concept art, thumbnails, character stills, storyboards, or ad mockups. The open question is whether a model can animate those assets quickly enough to make iteration normal rather than painful. Grok Imagine 1.5 is aimed directly at that use case: give it a starting image, describe the motion, and choose resolution and duration 1.
The new workspace layer points in the same direction. Projects make outputs easier to organize; multiple agents allow parallel generations; search reduces the archive problem after dozens of attempts 1. Those are not model-quality claims. They are signs that xAI wants Grok Imagine to absorb more of the editing-room routine around the model.
The limitation: this is not yet a clean enterprise story
The launch still leaves several questions open. xAI's public announcement does not provide a full pricing table for 1.5 video generation, and the API landing page is more explicit about image pricing than video pricing 3. The official post also emphasizes 720p Fast output, which is useful for social and draft workflows but below the 1080p or 4K claims that some competitors use to court professional production teams 1.
There is a second caveat: general availability is not the same as operational trust. Developers still need to test queue behavior, failure rates, content-policy behavior, audio consistency, prompt adherence, and whether the model's quality holds across less cinematic inputs than the showcase examples. Leaderboard position gets Grok into the evaluation set; it does not settle procurement.
Bottom line
Grok Imagine Video 1.5 is a significant release because it combines three signals in one move: a named GA API model, a faster consumer path, and a workspace layer that supports iterative production. It does not knock Seedance off the image-to-video-with-audio leaderboard, and the text-to-video story is less strong than the headline might imply. But xAI now has a credible video-generation surface that creators and developers can actually build around, rather than merely watch in benchmark screenshots.
References
- 1
- 2
- 3
- 4Artificial Analysis Image to Video Leaderboard
artificialanalysis.ai
- 5Artificial Analysis Text to Video Leaderboard
artificialanalysis.ai

Video Gen Model Tracker
An event-triggered channel covering major milestones in the video generation AI space. Every time Seedance, Kling, Veo, HappyHorse, or a notable competitor drops something significant — new model version, benchmark result, key feature — a dedicated article goes out with full context and analysis.
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
- Sign in to comment.