This follows Why Unreal Apps Wait on One Core. Read that first: it explains who Sam is, the one processor core at the kitchen door that prepares every picture.
Memory only stores the game world. Nothing happens to a rock just because it is in memory. The work starts when a picture has to be prepared, and that work grows with how many separate objects there are.
Back in the kitchen: before Sam passes an order to the cooks, Sam checks every ingredient it needs. One big sack of flour is one check, however heavy it is. Ten thousand peppercorns, each checked one by one, are ten thousand checks, however little they weigh. And Sam does this for every order, 35 times a second.
Unreal is the same. One huge, detailed texture is one check. Ten thousand pebbles, each a separate object, are ten thousand checks for every picture. Memory is the weight. Sam's work is the counting.
| in the game world | memory | processor | graphics card |
|---|---|---|---|
| One very detailed 4K texture about 20 MB |
a lot | almost none | little |
| 10,000 small rocks, each a separate object | little | a lot | little |
| The same 10,000 rocks grouped as one object questions asked once for the whole group |
little | little | little |
| One statue with millions of tiny details, using Unreal's Nanite one object to the processor; the graphics card handles the detail |
some | almost none | some |
| The 10,000 separate rocks with ray tracing on realistic reflections: every object adds more work per picture |
some | even more | some |
| 200 lights that cast shadows | some | a lot | a lot |
| 500 characters walking around | some | a lot | some |
| Showing the picture in 4K instead of HD | a little | none | 4× the pixels |
| Higher lighting and image quality settings | a little | none | a lot |
The short rule: memory pays for how big things are, the processor for how many separate things there are, and the graphics card for how many pixels it draws and how good each one looks.
Eagle 3D Streaming ran one large Unreal city scene on two cloud machines. Both have a strong graphics card. The difference is the speed of a single processor core.
| machine | one core runs at | price per hour | pictures per second |
|---|---|---|---|
| Cloud machine A | 3.4 GHz | $3.74 | 26 |
| Cloud machine B | 3.9 GHz | $3.78 | 35 |
On both machines, preparing the picture took longer than drawing it. The graphics card finished early and waited. On the machine with the faster core, preparation got quicker and the app went from 26 to 35 pictures a second, while the drawing time hardly changed.
0.5 GHz looks small on paper. On screen it was 9 more pictures every second, about a third more. The GHz number also undersells the difference: machine B's core is a newer design that gets more done in every tick, so Sam's preparation time fell from 36 to about 27 ms per picture. Two cores with similar GHz can still be very different Sams; the only sure way to know is to measure the app on the machine.
And the price hardly differed: about 4 cents an hour (AWS on-demand with Windows, N. Virginia, October 2026). The price of a cloud machine mostly follows its graphics card and memory, not how fast one of its cores is. A more expensive machine is not a faster one if the app is waiting for Sam.
Apps are usually built on a powerful desktop computer. When it is time to put the app in the cloud, the natural step is to look for a machine with the same graphics card. But the processor matters too, and here desktops and cloud servers are very different:
So an app that ran at 40 pictures a second on the developer's desk can run at 26 on a cloud machine with an equally strong graphics card. Nothing is wrong with the graphics card. The person at the kitchen door is simply slower.
The lesson: when choosing where an app will run, write down the developer's processor and its speed next to the graphics card, and compare like with like. Then test on the machine the app will actually run on.
| where | change | why |
|---|---|---|
| The machine | Choose by the speed of one core as well as by graphics card | A faster person at the door. |
| The app | Group many small objects into one (instancing, merging) | Fewer questions per picture. |
| The app | Use Nanite for detailed objects | Detail moves to the graphics card; the object counts once. |
| The app | Use ray tracing only on objects that need it | Removes extra work for every small object. |
| The app | Fewer lights that cast shadows | Fewer questions per picture. |
| The app | Test the final release build, not the development build | The development build does extra checking on the processor. |
The "person at the door" is Unreal's render thread. Each frame passes through four lanes in parallel: the game
thread (gameplay, animation, physics), the render thread (visibility, LOD, draw lists, ray-tracing scene update), the RHI
thread (DirectX 12 commands) and the GPU. The frame time is the slowest lane. In a running build, the console command
stat unit shows them as Game, Draw, RHIT and GPU: if Draw is larger than GPU, the render thread
is the limit and a faster GPU will not help. stat scenerendering shows draw calls and primitives; an Unreal
Insights capture (-trace=cpu,gpu,frame) shows where the render thread spends its time.
| machine | game | render | RHI | GPU | fps |
|---|---|---|---|---|---|
| AWS g6e.4xlarge, L40S, EPYC 3.4 GHz | 16 ms | 36 ms | 26 ms | 23 ms | ~26 |
| AWS g7.4xlarge, RTX PRO 4500, Intel 3.9 GHz | 12 ms | 26–29 ms | 16 ms | 20–22 ms | 34–37 |
Example measurements by Eagle 3D Streaming, October 2026: one Unreal Engine 5 scene at 1920×1080 on two AWS machines. Streamed and non-streamed runs gave the same frame rate, so streaming itself cost nothing measurable.
Last updated