isn't 11min a lot longer than most AI content made now?
A lot of AI content is a single prompt output that produces one continuous piece. This type of content is relatively short due to the limits of the context window.
Agentic models use a series of prompts to generate and patch together pieces using 'skills' (interfaces for the model to act in a determinative manner). There is also the possibility of human involvement in editing. There's basically no length limit aside from what you're willing to spend on compute for models of this kind, you can add more agent layers to create and combine multiple scenes and keep each scene within one context window's limits.
As far as the actual video, FF8 was my favorite offline FF and I still only made it like 30 seconds. Repeating the same dialogue and the flat tones and mismatched lip movements, ugh. That's not to say the tech can't eventually reach a point where it's viable.
My guess is a transition toward a model where the AI creates 3D assets and builds an intermediary format that controls the camera, placement, and tunables like lip movement/dialogue, then another tool is used to render the final scene. A model like that would allow much easier fine tuning and localization than the current method (which is mostly just repeating and praying your output is usable enough to patch together).