What slows Sam down

This follows Why Unreal Apps Wait on One Core. Read that first: it explains who Sam is, the one processor core at the kitchen door that prepares every picture.

It is the number of objects, not their size

Memory only stores the game world. Nothing happens to a rock just because it is in memory. The work starts when a picture has to be prepared, and that work grows with how many separate objects there are.

Back in the kitchen: before Sam passes an order to the cooks, Sam checks every ingredient it needs. One big sack of flour is one check, however heavy it is. Ten thousand peppercorns, each checked one by one, are ten thousand checks, however little they weigh. And Sam does this for every order, 35 times a second.

Unreal is the same. One huge, detailed texture is one check. Ten thousand pebbles, each a separate object, are ten thousand checks for every picture. Memory is the weight. Sam's work is the counting.

in the game worldmemoryprocessorgraphics card
One very detailed 4K texture
about 20 MB
a lotalmost nonelittle
10,000 small rocks, each a separate object littlea lotlittle
The same 10,000 rocks grouped as one object
questions asked once for the whole group
littlelittlelittle
One statue with millions of tiny details, using Unreal's Nanite
one object to the processor; the graphics card handles the detail
somealmost nonesome
The 10,000 separate rocks with ray tracing on
realistic reflections: every object adds more work per picture
someeven moresome
200 lights that cast shadows somea lota lot
500 characters walking around somea lotsome
Showing the picture in 4K instead of HD a littlenone4× the pixels
Higher lighting and image quality settings a littlenonea lot

The short rule: memory pays for how big things are, the processor for how many separate things there are, and the graphics card for how many pixels it draws and how good each one looks.

A real example: same kind of graphics card, different processor

Eagle 3D Streaming ran one large Unreal city scene on two cloud machines. Both have a strong graphics card. The difference is the speed of a single processor core.

preparing the picture (one processor core)drawing the picture (graphics card)
Time per picture in milliseconds (thousandths of a second). Shorter is faster. For 30 pictures a second, each one must be ready in 33 ms.
machineone core runs atprice per hourpictures per second
Cloud machine A3.4 GHz$3.7426
Cloud machine B3.9 GHz$3.7835

On both machines, preparing the picture took longer than drawing it. The graphics card finished early and waited. On the machine with the faster core, preparation got quicker and the app went from 26 to 35 pictures a second, while the drawing time hardly changed.

0.5 GHz looks small on paper. On screen it was 9 more pictures every second, about a third more. The GHz number also undersells the difference: machine B's core is a newer design that gets more done in every tick, so Sam's preparation time fell from 36 to about 27 ms per picture. Two cores with similar GHz can still be very different Sams; the only sure way to know is to measure the app on the machine.

And the price hardly differed: about 4 cents an hour (AWS on-demand with Windows, N. Virginia, October 2026). The price of a cloud machine mostly follows its graphics card and memory, not how fast one of its cores is. A more expensive machine is not a faster one if the app is waiting for Sam.

Same GPU, slower core: why the developer's computer misleads

Apps are usually built on a powerful desktop computer. When it is time to put the app in the cloud, the natural step is to look for a machine with the same graphics card. But the processor matters too, and here desktops and cloud servers are very different:

  • A desktop processor commonly runs one core at 5 GHz or more. It is built to do one thing very fast.
  • The cloud server processors Eagle 3D Streaming measured run at 3.4 to 3.9 GHz. They are built to do many things at once, not one thing fast.

So an app that ran at 40 pictures a second on the developer's desk can run at 26 on a cloud machine with an equally strong graphics card. Nothing is wrong with the graphics card. The person at the kitchen door is simply slower.

The lesson: when choosing where an app will run, write down the developer's processor and its speed next to the graphics card, and compare like with like. Then test on the machine the app will actually run on.

Three things that sound right but are not

"The processor is only at 59%, so the processor is not the problem." The usual CPU percentage is an average over all cores. What matters is the one core Sam is working on. If that core is at its limit, every picture waits for it, however idle the other cores are. In the example, machine B showed 59% overall, yet Sam's core was still the slowest part of every picture. Look at how busy Sam's core is, not at the average.
"Get a machine with more cores." More cores means more back-office staff. The person at the door is still one person. 32 cores do not prepare one picture faster than 16; a faster core does.
"Turn the graphics down." Lower resolution or lower image quality makes the graphics card's job easier. If the graphics card was already waiting, the app looks worse and runs no faster.

What actually helps

wherechangewhy
The machineChoose by the speed of one core as well as by graphics cardA faster person at the door.
The appGroup many small objects into one (instancing, merging)Fewer questions per picture.
The appUse Nanite for detailed objectsDetail moves to the graphics card; the object counts once.
The appUse ray tracing only on objects that need itRemoves extra work for every small object.
The appFewer lights that cast shadowsFewer questions per picture.
The appTest the final release build, not the development buildThe development build does extra checking on the processor.
For your developers: the same thing in Unreal's terms

Threads and how to measure

The "person at the door" is Unreal's render thread. Each frame passes through four lanes in parallel: the game thread (gameplay, animation, physics), the render thread (visibility, LOD, draw lists, ray-tracing scene update), the RHI thread (DirectX 12 commands) and the GPU. The frame time is the slowest lane. In a running build, the console command stat unit shows them as Game, Draw, RHIT and GPU: if Draw is larger than GPU, the render thread is the limit and a faster GPU will not help. stat scenerendering shows draw calls and primitives; an Unreal Insights capture (-trace=cpu,gpu,frame) shows where the render thread spends its time.

The example, lane by lane (CSV profiler, 1920×1080, Windows Server 2022, SM6, hardware ray tracing)

machinegamerenderRHIGPUfps
AWS g6e.4xlarge, L40S, EPYC 3.4 GHz16 ms36 ms26 ms23 ms~26
AWS g7.4xlarge, RTX PRO 4500, Intel 3.9 GHz12 ms26–29 ms16 ms20–22 ms34–37

Which features cost the render thread

Example measurements by Eagle 3D Streaming, October 2026: one Unreal Engine 5 scene at 1920×1080 on two AWS machines. Streamed and non-streamed runs gave the same frame rate, so streaming itself cost nothing measurable.

Last updated