Visko (Visko Platform Inc.), an AI startup based in California, opened public access to its video-generation foundation model "Orbis" on September 1 (local time). Instead of producing a clip of a few seconds and stopping, as conventional video AI does, Orbis keeps generating 4K video at 24 frames per second in real time, and when you change the instructions midway, the content shifts without the stream ever stopping. The company also disclosed that it has raised 10 million USD (about 1.6 billion yen) in a pre-seed round led by Llama Ventures.
※1 USD = 159 JPY (as of September 2, 2026)
A model that "runs" rather than "renders"
Most video-generation AI to date has worked the same way: type a prompt, wait, and receive a single finished clip. Once it is done, it is fixed. You cannot extend it or change it partway through.
Visko places its models in a new category it calls "Live Models." In the company's own words, these models "run" rather than "render." Orbis generates the first frame, streams it immediately, then generates the next frame conditioned on everything that came before, repeating this indefinitely. Founder and CEO Qing (Will) Yin compares the mechanism to how a large language model predicts the next token one at a time.
Because of this design, swapping the prompt during generation does not pause or restart the video; the world simply changes in place. Orbis supports three modes: text-to-video, image-to-video, and continuation of an existing video, and it accepts prompts in multiple languages.
The selling point: it holds together at hour scale
The biggest problem for real-time generation models has been that the longer they run, the more the footage falls apart. Yin puts it this way: every real-time system before this hit the same wall, where the world degrades the longer it runs, and Orbis was built to get past that wall.
According to the company's technical report, Orbis sustains hour-scale generation with no evident loss of quality or color drift while maintaining real-time 4K at 24fps. Two components are key. One is a multi-scale memory that preserves subjects, scenes, and style across chunks. The other is a latent world model that scores candidate "next moments" during inference and steers the output toward physically plausible motion. By grounding object movement and dynamics in this world model, the goal is video that is natural in motion, not just visually continuous.
On evaluation, Visko says Orbis leads the real-time and long-video systems it was compared against on DOVER, which measures aesthetic and technical quality, and on VideoAlign, which measures visual and motion quality. In a human preference study of long-form video across eight systems, Orbis received the highest ratings for overall preference and temporal stability. The report also discloses dimensions where Orbis trails other systems, and the full methodology and results are published. The technical report was also posted to arXiv on July 29.
A 16-person team led by a former Apple researcher
Visko was founded in 2025 and is headquartered in Sunnyvale, California. Founder Yin holds a Ph.D. in computational mathematics and mechanics from Stanford University and spent three years as a researcher at Apple. The 16-person team draws from Apple, Google DeepMind, Meta, Amazon, and Tesla.
The advisory lineup stands out as well. The chair of the advisory board is Michael I. Jordan of the University of California, Berkeley, a researcher widely regarded as a leading figure in machine learning. He is joined by Steve WaiChing Sun of Columbia University and Mengye Ren of New York University. Jordan commented that Orbis has a meaningful technological lead over conventional short-clip video generation and that the same core capability could support a broad range of commercial opportunities, from robotics and physical simulation to gaming, live commerce, education, and real-time creative media. Distribution is handled by Visko's partner Reactor.
The target: training data for robots
The first use case Visko has in mind is robotics rather than entertainment. Yin points to generating footage of situations that are dangerous or expensive to stage in reality, for use in simulation and as training data. He also says Orbis can provide environments for testing rare situations (corner cases) to validate the safety of autonomous vehicles and humanoid robots.
At the same time, the company is candid about what is still missing. On the research side, the tasks are to further refine the memory structure to improve consistency across a generation, and to add fine-grained action descriptions such as "raise the arm 45 degrees" to the training data. On the engineering side, scaling up both the model and the data is next. Audio and video data are the current focus, and the company is exploring whether to incorporate sensor data such as touch. Yin expects that, as with autonomous driving, it will take time before general-purpose robots enter homes, and says the company intends to find commercialization opportunities along the way.
Summary
Visko's "Orbis" has been opened to the public as a "Live Model" that continuously generates 4K video at 24fps in real time and keeps going even when the prompt changes mid-stream. Its main claim is that quality and color hold up at hour scale, and in human evaluation it claims the top spot for overall preference and temporal stability among eight systems. With 10 million USD raised, the 16-person team led by a former Apple researcher is aiming less at generating video content and more at a "world that keeps running" for training robots and autonomous vehicles.
