On September 8, OpenAI published "The Work Now Within Reach," a document setting out how its business fits together. Alongside more than 1 billion weekly users and 2.5 million business customers, it puts numbers to how usage deepens the longer someone has an account, and explains where its own inference chip sits in the effort to cut serving costs[1].
Usage grows over the first six months
The most striking figure comes from a study of people on individual ChatGPT plans. Six months after signing up, daily message volume was roughly 50 percent higher than in the first month, and users had tried roughly twice as many distinct tasks[1]. People start with a narrow set of uses and hand over more as they get comfortable, and that pattern now shows up in the data.
OpenAI has built its revenue structure around that behavior. Free access supported by advertising serves as the place to discover where AI is useful, and people who find value move to subscriptions or usage-based plans. The point is to have the container ready before usage expands[1].
The scale numbers were disclosed as well. Across ChatGPT, ChatGPT Work, Codex and applications built on the API, weekly active users exceed 1 billion and business customers number 2.5 million. A single research advance improves several products at once, which spreads revenue across multiple streams[1].
The line between home and work is thinning
The document also describes how individual use and enterprise deployment pull on each other. Someone comfortable with ChatGPT at home brings that expectation to the office, and experience with complex work then changes what they expect from AI in their personal life. Developers widen the surface further by building applications for needs OpenAI would not have identified on its own. As agentic products get to know users as people, the company expects those two segments to blur continuously[1].
Its own research organization serves as the in-house example. For every workday of human labor, agents now contribute 3.1 agent-workdays of effort, and infrastructure problems that once required specialist support are being handed to agents. People still set priorities and judge results; the gain is in how many promising ideas a team can pursue[1].
Software and silicon aimed at inference cost
If demand keeps growing, the next lever is how cheaply and quickly a single response can be delivered. OpenAI says it manages data centers, chips, software, models and products together as a full-stack strategy, choosing the best combination for each workload[1].
There are concrete figures. Using GPT-5.6 Sol to improve its production serving software cut end-to-end serving costs by 20 percent, and further improvements raised token-generation efficiency by more than 15 percent[1].
On the hardware side, that work extends into Jalapeño, the company's first custom inference chip. In InferenceX tests across three public models, it delivered 1.5 to 1.9 times as much peak token throughput per watt as the commercial systems tested, using rated chip power to normalize the comparison, with end-to-end latency 1.7 to 3.6 times lower. Deployment is planned to begin by year-end alongside accelerators from NVIDIA, AMD and other partners[1].
Better models reduce the number of attempts needed; better software and hardware make each attempt faster and cheaper. Pushing on both, OpenAI argues, means serving more work from the same capacity[1].
Investment judged by how fast it pays back
The document also names the brake on that expansion. Each investment is judged by the demand it can serve, how quickly it becomes productive, and whether the returns justify the capital committed[1]. In an industry pouring enormous sums into compute, writing down the yardstick is itself a signal.
On the model side, GPT-6 Astra is positioned as state of the art in computer use, browsing, software engineering, cybersecurity, science and professional work[1].
Summary
This was not a product launch but a statement of how OpenAI describes its own loop. A 50 percent rise in daily messages and roughly double the range of tasks after six months suggests that once AI use takes hold, it thickens on its own. Absorbing that growth means cutting inference costs, and the savings buy more usage in turn. The individual figures matter less than how fast that loop can be turned, which is where the competition now sits.
Source: https://openai.com/index/the-work-now-within-reach/
* The thumbnail image is AI-generated and provided for illustration only.
