The Second Wave of Fake AI GTA 6 Gameplay Videos Already Has 1 Million Views (and Nobody's Going to Apologize)
There's a "leaked GTA 6 gameplay" making the rounds right now with over a million views. It doesn't exist. It's not a dev build, it's not internal capture, it's not anything. It was generated by AI, from the first frame to the last.
And it's not the first time. This is the second wave, and it's bigger than the 2025 one.
Instead of just pointing fingers at the fake, let's do the thing that matters to people who build: tear the pipeline apart. Because how GTA 6 AI videos are made is a free masterclass in consistent video generation. Better yet, every flaw that gives the video away is a direct consequence of a technical decision made by whoever produced it. You learn how to make them and how to spot them in the same read.
TL;DR
- What it is: a wave of fake AI-generated GTA 6 gameplay videos, with millions of views and DMCA takedowns from Take-Two.
- Stack in use: Veo 3.1 and Kling for video, voice cloning for audio, and reused official audio in some cases.
- A piece off the board: Sora 2 is out of the game. App shut down in April 2026, API shuts off on September 24, 2026.
- Why now: Rockstar announced an "Extended Look" for August 27. Until then, the information vacuum is fuel.
What's happening right now
The math is simple and brutal: Rockstar stays silent, the anticipation builds, and the vacuum gets filled by whoever has a video generator and a social media account.
A video from the Nexifygaming channel passed a million views in five weeks on Instagram, according to ComicBook's reporting. Others racked up millions more. Take-Two responds with DMCA takedowns, but the removal always shows up after the audience does.
In Brazil the story hit the news cycle on August 4, with Voxel publishing a piece about AI-made GTA 6 gameplay videos taking over the internet. The next day, an account went viral for remaking the entire trailer with AI, keeping the original audio on top.
And there's a clock: Rockstar confirmed an "Extended Look" at the game for August 27, premiering on Netflix six hours before YouTube. Translation: there isn't much vacuum left, and the vacuum is exactly where this kind of content thrives.
The first wave was in 2025, and the creator apologized
Worth putting on the record because it changes how you read what's happening now.
In November 2025, the ZapActu account posted a "leaked gameplay" that pulled 8 million views in 24 hours before it was exposed and deleted. The interesting part is what came next: they publicly admitted it was AI-generated, said they had no financial motive and only wanted to "observe people's reactions" and show how easy it had become to produce plausible content with AI. They promised to start labeling speculative content. Kotaku and CBC covered the episode.
That was an experiment with an apology at the end.
The 2026 wave has no apology. It has channels, a publishing cadence, and volume. What used to be a proof of concept became an audience operation. That's the difference that matters: it's not the technique that evolved the most, it's the intent.
How GTA 6 AI videos are made: the 5-step pipeline
I reconstructed the pipeline by cross-referencing what the creators themselves publish, the artifacts the press pointed out, and the known behavior of the models. Five steps.
1. Reuse the audio that already exists
The most underrated step, and the most effective. Some of these videos don't generate any audio at all: they take the score and narration from the official trailers and regenerate only the picture on top.
That solves, for free, the hardest problem in video generation, which is sync. The viewer's brain uses audio as a temporal anchor. If the audio is legit and well produced, it "carries" images that couldn't sustain the illusion on their own.
2. Prompt in JSON, not prose
This is the part that matters if you work with this stuff. Current video models follow structured instructions far more faithfully than running text. A prose description becomes an average; JSON becomes a spec.
The structure in use looks roughly like this, applied here to a neutral scene instead of someone else's IP:
{
"shot": "plano baixo seguindo um sedã pela avenida costeira",
"camera": "câmera na altura do para-choque, movimento lateral lento",
"world": "cidade litorânea noturna, art déco, palmeiras, letreiros de neon",
"grade": "magenta e ciano saturados, brilho úmido no asfalto",
"render_style": "captura de jogo em terceira pessoa, HUD ausente",
"duration_s": 8,
"negative": "sem texto na tela, sem logotipos, sem marca d'água"
}
Every field does a job. world and grade lock the visual identity across separate generations, which is what gives continuity to a video assembled from a bunch of short clips. render_style is what makes the model return something that looks like a game, not a film. And a short duration_s isn't an aesthetic choice: it's a limitation. The models hold coherence for a few seconds, so the final video is a patchwork of short clips stitched together.
3. One shot per generation
Nobody asks for an entire complex scene. They ask for one specific shot, generate it, throw away what came out bad, generate again. The discard rate is high, and that's where the real work lives. What lands in your feed is the survivor of dozens of attempts.
4. Cloned voice for whatever the original audio doesn't cover
When new dialogue is needed, voice cloning comes in. And this is where the economics of producing at scale show up: cloning a voice well takes work, so they clone only a few.
5. Distribution that mimics official marketing
Titles in the "Final Trailer (2026) Rockstar Games" pattern, thumbnails in the style of the campaign, posting at peak hours. Kotaku documented that YouTube itself went as far as sending a push notification recommending one of these fake videos. The algorithm doesn't tell the difference; it measures retention, and retention is something the fake has plenty of.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãSora left the stage in the middle of the wave
A detail that redraws the tooling map and that almost nobody connected.
OpenAI discontinued Sora. According to the official deprecations page, developers were notified on March 24, 2026, and the Videos API along with the sora-2 and sora-2-pro models shut off on September 24, 2026. The web experience and the app were shut down before that, in April.
In other words: the tool that became synonymous with AI video in 2025 isn't the one producing the 2026 wave. The work migrated to Veo 3.1 and Kling. If you maintain any product leaning on the Videos API, that date matters more to you than any GTA fake.
It's worth comparing with what we already tested here in Gemini Omni Flash in practice: the bottleneck was never generating a pretty clip. It's keeping coherence between clips.
The artifacts are the pipeline leaking
This is the part that only makes sense once you understand the five steps. The signs that give the video away aren't random. They're the pipeline showing through.
The same voice on two different characters. ComicBook pointed this out in one of the viral videos. It's not a model bug: it's step 4. They cloned a few voices and recycled them.
Wrong text on screen. Signs, subtitles, and HUD with weird spelling. And the most ironic part: the creators instruct the model not to generate text, like the negative field in the JSON up there. They do it precisely because the models still write badly. When text slips through, it slips through crooked.
Cuts that are too short. If no shot lasts longer than eight or ten seconds without a cut, that's not an editing choice. It's step 2: the model's coherence window ran out.
Physics that "slips." A car that changes proportions between shots, a shadow that doesn't line up, a reflection that doesn't match the scene. Per-clip generation has no global world state.
Where this stops being technique and becomes a problem
Reproducing this pipeline on top of someone else's intellectual property and monetizing it is not a gray area. Take-Two is issuing DMCA takedowns and has every right to.
The technique is neutral and legitimately useful: character consistency, continuity between clips, camera direction by spec. That works for product previews, campaign assets, animated storyboards, training material. What changes the game is the subject and the label. Generating a coastal city at night is work. Generating Vice City and calling it a leak is something else.
And there's an indirect cost that anyone producing content should weigh: every one of these waves burns a little of the public's default trust in video. Before long, legitimate material will have to prove it's real too.
Quick FAQ
Can you know for sure whether a video was AI-generated? With absolute certainty, no. You can add up the evidence: a voice repeated across characters, crooked text, very short cuts, inconsistent physics between shots. On its own, each one is weak. Together, they're pretty conclusive.
Does GTA 6 use generative AI in development? According to Take-Two's CEO, no. He told CNBC that generative AI has zero involvement in the game, which is built building by building. The irony of the story: the AI is all around the game, not inside it.
Is this JSON prompting technique worth learning? Yes, and it has nothing to do with GTA. Specifying a scene as structure instead of prose is, today, the difference between a reproducible result and a lottery. It transfers to any video model.
Does Sora still work? Until September 24, 2026 on the API, priced per second of generation. After that it shuts off, and OpenAI hasn't announced a replacement.
What's left
The story here isn't "AI makes fake video." We already knew that.
The story is that a pipeline assembled from off-the-shelf tools, running on top of a company's communication vacuum, delivers an audience in the millions before any legal team can react. The speed asymmetry between producing and taking down is the real problem.
On August 27, Rockstar shows official material. My bet is the flood of fakes grows instead of shrinking, because fresh official material is raw material for generating more convincing fakes.
And the skill that's left over from this story, for people who build: knowing how to spec a scene in JSON, understanding where coherence breaks, and recognizing the artifact before you share. This isn't about GTA. It's about working with generative video without being a passenger on it.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã