A small Reddit thread has surfaced a surprisingly good strategic question for OpenAI. Now that Sam Altman has publicly explained that the company killed Sora partly because video generation consumed compute it wanted to devote to Codex, could OpenAI’s new custom inference chip eventually change the economics enough to bring advanced video generation back — not as another standalone social app, but as a capability inside ChatGPT?
The r/ChatGPT post comes from a creator who says they generated thousands of pieces with Sora and misses both the video tool and the community that grew around it. Their proposal is speculative: once OpenAI deploys its Jalapeño inference hardware at scale, perhaps a future “Sora 3” could return inside ChatGPT under a strict credit system rather than operating as a compute-hungry standalone product.
There is no announced Sora 3, and OpenAI has given no indication that it plans to resurrect the consumer Sora product. But the underlying question is worth taking seriously because two things have changed. Altman has now been unusually explicit about why Sora lost OpenAI’s internal resource battle, and OpenAI has published the first performance results from hardware designed specifically to make inference cheaper and faster. Together they expose the real constraint behind many AI products: a feature can be technically impressive, popular and strategically interesting and still die because another workload produces more value from the same unit of compute.
Altman says Sora was killed because something else mattered more
OpenAI did not shut Sora because Altman thought it was a bad product. In an August interview with David Senra, the OpenAI CEO described Sora as good, fun and cool, then explained that it used a large amount of compute and was less important than Codex. He made a similar point about the Atlas browser, which OpenAI abandoned despite Altman praising the product, because the company needed to concentrate limited people and resources elsewhere.
The comments clarify a strategic shift that had already become visible earlier in 2026. OpenAI is increasingly organizing itself around frontier models, coding and knowledge-work agents, enterprise customers and the infrastructure required to serve those systems at enormous scale. Consumer experiments are being judged against that opportunity cost.
That is a harsher test than simple product-market fit. Sora had users. It demonstrated technology that only a few years earlier would have looked extraordinary. Its social app attempted to build a new kind of creative network around generated video. Yet every expensive video generation also competed for infrastructure that could instead serve coding agents, business workflows and general ChatGPT demand.
Altman’s explanation effectively turns compute into a capital-allocation problem. A GPU-hour is not just a technical resource. It has an alternative economic use.
Video has particularly unforgiving inference economics
Text-to-video is one of the clearest examples of why AI product economics cannot be inferred from subscription prices alone. A language model generates sequences of tokens. A video system has to produce a coherent visual world across many frames, maintain identities and motion over time and, in newer systems, potentially generate synchronized audio. Users also commonly create several attempts before keeping one result.
That means a consumer can discard a ten-second clip after watching it once while the provider still absorbs the full inference cost of creating it. A social feed makes the economics even more demanding because it encourages frequent generation and experimentation rather than occasional professional use.
Reporting around Sora’s shutdown put its infrastructure cost at roughly $1 million per day during part of its operation. Exact product-level economics remain private, and the figure should not be mistaken for a permanent fixed cost. But it illustrates the scale of the problem: consumer video can burn substantial compute before the company has found a business model that captures equivalent value.
Codex has a different economic profile. If an AI coding agent saves an engineer hours, fixes a production problem or completes a business workflow, a company can compare the cost of the AI directly with labor and project value. The willingness to pay for successful knowledge work can therefore be much higher than the willingness to pay for another experimental social video.
Jalapeño changes the equation — but not in the simplistic way Reddit hopes
The Reddit argument becomes interesting because OpenAI has now demonstrated its first custom inference chip. In its August 25 Jalapeño benchmark announcement, OpenAI said the accelerator delivered between 1.5 and 1.9 times more AI work per watt at peak throughput and between 1.7 and 3.6 times lower end-to-end latency than the comparison systems across several public large language models.
OpenAI designed Jalapeño with Broadcom specifically around LLM inference. Its original June announcement described the chip as the first generation of a broader OpenAI compute platform, with deployment beginning in 2026 and expanding over multiple generations. The company says it intends to optimize chips, models, kernels, networking, memory and serving software as one integrated stack.
If OpenAI can serve more useful AI work from the same amount of electrical power and infrastructure, the consequences can be enormous at its scale. Lower inference cost can support higher usage limits, more agent steps, lower API pricing or products that were previously too expensive to operate broadly.
But there is an important technical caveat to the “Jalapeño brings Sora back” theory. Jalapeño is an inference accelerator built around LLM workloads. OpenAI has not said it was designed specifically to serve Sora-style diffusion or video-generation architectures, and inference is only one part of the video economics. Training the next frontier video model remains enormously expensive. So do storage, safety systems, moderation and media delivery.
A better chip can improve the cost curve without magically making every abandoned product rational.
The more plausible comeback is video inside ChatGPT, not Sora as another app
The Reddit user’s strongest idea is therefore not the name “Sora 3.” It is the proposed location: inside ChatGPT.
OpenAI’s recent product decisions suggest that the company wants fewer independent destinations and more capabilities concentrated around a unified interface. Altman has described OpenAI as a platform company rather than a company that needs to operate a separate consumer product for every modality. Sora and Atlas illustrate the difference. OpenAI can abandon a video social network without abandoning video research, just as it can abandon a standalone browser without abandoning agents that use the web.
ChatGPT already gives OpenAI the distribution layer. A future video-generation capability would not necessarily need its own feed, social graph, discovery system, mobile application, creator community infrastructure and separate product team. It could appear as another tool invoked when the conversation requires moving images.
That architecture is potentially more efficient organizationally as well as computationally. The user could storyboard in chat, generate still references, iterate on characters, create a video and revise it without leaving the same workspace. The model could preserve project context instead of forcing the creator to move prompts between separate products.
In that sense, killing Sora may eventually look less like abandoning AI video and more like rejecting one particular packaging of AI video.
A credit system would make more sense than unlimited generation
The Reddit proposal also points toward the most obvious economic mechanism: credits. Expensive generative workloads do not fit comfortably inside subscriptions marketed as broadly unlimited access. A credit system lets the provider expose the capability while making its scarcity visible.
A ChatGPT subscriber could receive a monthly allocation of video credits, with different resolutions, durations and quality levels consuming different amounts. Heavy creators could buy additional capacity. Professional plans could include larger quotas and commercial workflow features. Free users might receive occasional low-cost generations or sponsored access.
This is less psychologically attractive than the early era of AI products, when companies often subsidized expensive generation to accelerate adoption. It is probably more sustainable. Video generation has a physical cost in accelerators, electricity, cooling and data-center capacity. If the workload remains expensive even after hardware improvements, product design eventually has to expose that fact.
OpenAI’s decision to introduce advertising into ChatGPT also shows that the company is willing to experiment with multiple ways to subsidize consumer inference. Ads are currently intended to support Free and Go access to ChatGPT rather than fund a hypothetical video service, but the broader business logic is relevant: the company is searching for revenue structures that let more users consume costly intelligence without every interaction being covered by a high subscription fee.
The chip may matter more for Codex than for Sora
There is another reason to be cautious about predicting a video revival. Jalapeño’s published strengths align extremely well with the product OpenAI chose over Sora.
Agentic coding is highly sensitive to latency because an agent may perform dozens or hundreds of sequential reasoning and tool-use steps. A small delay repeated across every step becomes a substantial delay in the overall task. OpenAI explicitly highlights agents when explaining why Jalapeño’s lower latency matters.
That means the custom chip can reinforce the Codex strategy rather than reverse it. If OpenAI can make coding agents faster and cheaper, the return on allocating additional compute to Codex may increase too. Sora does not merely need to become cheaper; it needs to become competitive with an alternative workload whose economics are improving at the same time.
This is the central mistake in thinking about compute efficiency as if it automatically resurrects abandoned products. When the cost of intelligence falls, every AI application benefits. The scarce resource may remain scarce because demand expands faster than efficiency.
Jevons paradox may arrive for AI compute
Computing has repeatedly demonstrated a pattern in which efficiency improvements do not reduce total resource consumption. They make new uses economical, which increases demand. AI is already showing signs of the same effect.
If Jalapeño lets OpenAI serve twice as much inference from a given power budget, the company does not necessarily end up with half its data center sitting idle. It can run more ChatGPT requests, let Codex agents take more steps, increase context, lower prices, support more users or deploy more capable models. Efficiency creates capacity, and product teams immediately find ways to consume it.
OpenAI’s own roadmap points in this direction. The company describes Jalapeño not as a way to reduce its need for infrastructure but as part of a multigenerational platform that will be deployed at enormous scale. Better economics make intelligence more abundant, but abundance stimulates new demand.
A Sora comeback would therefore depend on strategic priority, not simply spare chips appearing in a rack.
Video still has strategic value that Codex cannot replace
There is nevertheless a reason OpenAI may eventually want advanced video back in its flagship product. Coding agents dominate economically valuable knowledge work, but video generation offers something different: consumer creativity, media creation and potentially a route toward models that understand the physical world.
OpenAI did not abandon the underlying world-model research when it closed Sora’s consumer experiences. Research associated with Sora has continued to matter for physical reasoning, simulation and robotics. A model capable of generating coherent video has to learn aspects of objects, motion, space and temporal consistency that are relevant beyond entertainment.
That means future video research could return to users as a consequence of broader multimodal progress rather than because OpenAI decides to rebuild the old Sora social network.
The product might not even be called Sora. OpenAI’s current consolidation strategy makes modality-specific branding less important. Users increasingly expect one assistant to write, code, research, browse, generate images, analyze files and operate tools. Video can become another output mode of that assistant.
The community Sora lost is harder to reconstruct than the model
The Reddit post also highlights something infrastructure discussions tend to ignore: Sora was not only a model endpoint. For some creators it became a community. Users shared clips, prompts, techniques and visual experiments in a dedicated environment. That social layer generated inspiration and made the product feel like a destination rather than a utility.
Integrating video into ChatGPT could restore the capability while failing to restore that experience. A private assistant is structurally different from a creative network. ChatGPT can help one creator produce a film, but it does not automatically provide a feed where thousands of creators discover one another’s experiments.
This may have been one of the strategic costs OpenAI accepted when it chose focus. Social networks require sustained moderation, recommendation systems, creator incentives and community management. They can become powerful distribution assets, but they also turn the company into a media platform. OpenAI appears more interested today in becoming the intelligence layer beneath work and applications.
There is no Sora 3 — but the conditions for a different return are becoming clearer
It would be easy to turn the Reddit discussion into a rumor: OpenAI has a new chip, therefore Sora 3 is coming. There is no evidence for that conclusion. OpenAI has not announced Sora 3, and Altman’s recent comments explain why the company deliberately walked away from the consumer video product.
What has changed is the underlying economics. OpenAI now owns working inference silicon designed around its workloads. It is consolidating products around ChatGPT and Codex. It is developing multiple monetization systems, including subscriptions, enterprise contracts, APIs and advertising. And it continues to research multimodal and physical-world intelligence even after shutting down Sora’s standalone experience.
Those ingredients make a future return of advanced video plausible in a form that looks very different from Sora 2: integrated into ChatGPT, metered according to compute cost and connected to a broader multimodal workflow rather than built around a separate social app.
The key question is not whether Jalapeño makes video cheap. It is whether video becomes valuable enough, within OpenAI’s unified platform, to win the compute-allocation battle it already lost once.
Altman’s explanation of Sora’s death gives creators a frustrating answer, but also a useful one. OpenAI did not decide that generative video had no future. It decided that, at that moment, every unit of scarce infrastructure could do something more strategically important in Codex. If custom silicon and better models shift that calculation far enough, video may return. It just may not return as Sora.