普通视图

发现新文章,点击刷新页面。
昨天以前Tomshardware

We tested unofficial DLSS Multi Frame Generation support on RTX 40-series GPUs — new mod brings RTX 50-series exclusive feature to older cards, and it really works

2026年9月12日 22:08

It’s been a heck of a time lately for PC gamers willing to get their hands dirty with mods. Hot on the heels of the discovery of the DLSS 5 DLL in a prerelease version of NBA 2K27, modder dashdogy found a way to bring Multi Frame Generation, one of the crown jewels of GeForce RTX 50-series graphics cards, to RTX 40-series (and earlier) products.

As already elevated graphics card prices seem set to continue rising, and hardware upgrades get further and further out of reach of the average PC gamer, more and more folks are going to want to hold on to the RTX 40-series hardware they have for as long as they can, especially if smoothness-boosting features like MFG are just a few clicks away on those older cards.

So we had to see MFG working on Ada for ourselves—assuming it works at all. We grabbed the mod files and got to playing with them in Cyberpunk 2077, since it’s likely in many TH readers’ Steam libraries already and has a healthy modding community. You can find the latest instructions for enabling MFG on Ada through dashdogy’s GitHub page.

Before spending a ton of time testing, we verified that the mod works at all. While spinning the camera at a high, constant speed, we could indeed see that increasing MFG multipliers beyond the officially supported 2X factor on Ada cards does greatly increase perceived smoothness or fluidity of motion in Cyberpunk 2077, as you would expect.

Another tell is that the same visual artifacts are visible in certain regions of the screen on both RTX 40-series and RTX 50-series graphics cards as you add more generated frames. MFG 4X and above, especially, tend to add some visual “junk” at the bottom of the frame that appears regardless of the game, and we could see that artifacting on Ada.

At least in Cyberpunk 2077, then, we’re confident that MFG is really doing its thing on RTX 40-series cards with this mod.

We also didn’t see any perceptible issues with frame pacing or frame delivery, although playing on a high-refresh-rate, G-Sync-Compatible monitor like our test bench’s ROG Strix XG27UCS smooths out all but the worst such issues. If you don’t have a high-resolution or high-refresh-rate monitor to begin with, the utility of MFG will be seriously limited for you anyway.

The latency question

So is this a free lunch? Is Nvidia soft-locking MFG to Blackwell purely for marketing reasons? That’s where performance testing comes in, since it has the potential to reveal whether there’s a catch in running MFG on Ada. The question is not so much whether MFG juices output frame rates, but whether it runs on Ada within acceptable latency thresholds.

As we’ve long emphasized, when you have essentially arbitrary control over output frame rates like MFG allows, input lag becomes the final barrier to a playable experience. So in the performance results that follow, we’ll certainly present output frame rates as you would expect. But you should view those in the context of your own monitor’s refresh rate. As long as the delivered frame rates we recorded exceed your display’s peak refresh rate, you have headroom to play with to keep your monitor at or near that number during gameplay.

The real issue, then, is whether both RTX 40-series and 50-series cards deliver an acceptable input latency under our test conditions. In our past testing, we’ve determined that a roughly 60ms average latency threshold, as indicated by Nvidia’s FrameView app, is the point at which player inputs and displayed frames start to become noticeably decoupled in AAA single-player experiences like Cyberpunk.

If you’re only slightly on the wrong side of this threshold, a game might still be playable, but if you totally blow past it, you’re likely to notice laggy inputs and increasingly distracting visual artifacts as the MFG model struggles to fill in the gaps between sparser and sparser input data.

Testing methods and notes

As one of the biggest technical showcases of the current PC gaming era, Cyberpunk 2077 lets us enable all the modern rendering features we’d want for Nvidia cards. It implements not only ray tracing and path tracing, but DLSS Super Resolution, Ray Reconstruction, and Multi Frame Generation. The number of AI-generated pixels per frame can be quite high in this title.

We enabled all those features to expose the full potential complexity of running all of their associated AI models in a modern rendering pipeline. If RTX 40-series cards are going to stumble for some reason with modded MFG enabled, we want to put as many obstacles in their way as possible.

And because benchmarking the performance of an unofficial mod is venturing into the Wild West anyway, we also added a DLSS 5 mod to the mix to see whether Blackwell GPUs have a distinct edge in the neural rendering future that the arrival of that feature promises to usher in. DLSS 5 is supposed to come to RTX 40-series cards later this year, so we think it’s good to understand where performance sits today, even if it’s subject to the same disclaimers as this MFG mod.

For reference, then, we tested Cyberpunk with maxed-out raster settings, path tracing, and MFG 4X at three resolutions: 1080p with DLSS Balanced, 2560x1440 with DLSS Performance, and 4K with DLSS Ultra Performance upscaling enabled.

We only tested cards ranging from the RTX 4090 down to the RTX 4070 for these experiments, both because of time and Cyberpunk’s VRAM requirements. We didn’t want to deal with potential performance pitfalls due to running out of VRAM, so the 12GB RTX 4070 is where we’re drawing the line for now.

Modded Cyberpunk 2077 1080p performance

MFG on Ada

(Image credit: Tom's Hardware)

At 1080p, Cyberpunk 2077 with MFG 4X scales fine on Ada, but it’s clearly scaling better on Blackwell. The RTX 4090 lands between the RTX 5070 Ti and RTX 5080. Introducing DLSS 5 to the mix doesn’t change any relative standings, although it does create some larger gaps between average frame rates and 1% lows than we might like on the RTX 4070 Ti Super, RTX 4070 Ti, RTX 4070 Super, and RTX 4070. The RTX 5070 has no such trouble.

MFG on Ada

(Image credit: Tom's Hardware)

But again, the real story is in our latency results. Here, we can see that every card we tested falls under our 60ms threshold with MFG 4X alone. But Blackwell hardware has a clear advantage in the standings, as even the RTX 5070 delivers a lower input latency than the RTX 4090 with our DLSS 5 mod off and a comparable input latency with it enabled. All the other Ada cards shake out as you would expect from there. But in absolute terms, even with DLSS 5 enabled, only the RTX 4070 Super and RTX 4070 are far beyond our acceptable latency thresholds.

Blackwell might have a latency edge with MFG enabled and a smoothness edge with a modded version of DLSS 5 on top, but at least with the RTX 4070 on up, there’s certainly enough headroom to use the feature on Ada.

And as we went to press, a version of the RTX 40-series MFG mod came out that removes the need for ReShade. That more streamlined approach might cut down latency, but we couldn’t test it because it currently crashes Cyberpunk 2077. Again, this is the Wild West, not a validated, bulletproof solution from Nvidia like you get with RTX 50-series products.

Modded Cyberpunk 2077 2560x1440 performance

MFG on Ada

(Image credit: Tom's Hardware)

At 2560x1440, generational performance standings and scaling between cards remains much the same as we saw at 1080p. The RTX 4090 still falls short of the RTX 5080, and the RTX 4080 duo falls behind the RTX 5070 Ti.

We also still see the wide gap between average frame rates and 1% lows rear its head on more Ada cards with our DLSS 5 mod enabled, though to be fair, these lows are still being smoothed over by MFG to the point that you’re unlikely to notice them with a high-refresh-rate, variable-refresh-rate monitor like we’re using.

MFG on Ada

(Image credit: Tom's Hardware)

On the latency side, all of the Ada cards except the RTX 4070 still run our modded MFG 4X with acceptable input latency. But enable the DLSS 5 mod we’re using, and input latency climbs past 60ms for all Ada cards except the RTX 4090. The RTX 4080 Super and RTX 4080 still provide acceptable performance under this full load, as they’re only slightly over our latency threshold. But for any less powerful RTX 40-series cards, you’d need to start choosing between DLSS 5 and other eye candy, like lighter RT settings instead of path tracing or no RT or PT at all.

Modded Cyberpunk 2077 4K performance

MFG on Ada

(Image credit: Tom's Hardware)

Our modded performance results at 4K with DLSS Ultra Performance demonstrate why it’s so important to discuss MFG-boosted frame rates in the context of input latency. If you’re not mentally dividing by four, everything on our output frame rate might look playable.

At this high output resolution, the RTX 4090 finally takes the lead over the RTX 5080, both with plain MFG 4X and with our DLSS 5 mod on top. The 16GB of VRAM of the RTX 4070 Ti Super would seem to be giving it an edge over the RTX 4070 Ti and RTX 5070, and the RTX 4080 duo would seem to beat out the RTX 5070 Ti.

MFG on Ada

(Image credit: Tom's Hardware)

Our latency chart for MFG 4X shows a weird, but entirely reproducible result: input latencies at 4K with DLSS Ultra Performance are actually much lower than they are at 2560x1440 for most Ada cards, despite the fact that both of these output resolution targets share the same input resolution.

We’re not sure why Ada cards run into such a latency hump with this MFG mod at 2560x1440 with DLSS Performance, but we double-checked our results, and this behavior is reproducible. Blackwell cards experience the more linear rise in input latency as output resolutions rise that you would expect.

Again, this is the Wild West of modded performance, and we have nobody to blame but ourselves here, but it’s an unfortunate result given the prevalence of 2560x1440 monitors. Perhaps the maintainers of this mod can track down the root cause and fix it, but nothing is guaranteed.

In any event, input latency with MFG 4X alone at 4K with DLSS Ultra Performance isn’t an issue for any card here. You might be pushing your luck with the RTX 4070, but both Blackwell and Ada cards are delivering on the promise of MFG here: smoother output with responsive input.

Add DLSS 5 to the latency picture, though, and as we’ve come to expect, you really want an RTX 5090, RTX 4090, or RTX 5080 for acceptable responsiveness. And you’re pushing it with the RTX 5080. No other cards in this bunch need apply.

Bottom line

Our experience with the purportedly Blackwell-exclusive Multi Frame Generation on RTX 40-series cards through modding suggests that there isn’t any glaring reason why Nvidia couldn’t enable the feature for Ada Lovelace cards, and that’s kind of wild given how heavily it was touted as a Blackwell-exclusive feature back when those cards launched.

At least as long as Nvidia doesn’t find a persistent way to lock it out, our experience is that MFG generally just works on Ada. Even if input latencies aren’t quite as low on those older cards as they are on comparable Blackwell hardware, all else equal, they’re still perfectly acceptable, even under the combined load of path tracing, DLSS Super Resolution, DLSS Ray Reconstruction, and MFG 4X in Cyberpunk 2077.

We only had time to test cards ranging down to the RTX 4070 for this quick look, but for folks looking to extend the useful life of their Ada hardware, the availability of MFG could certainly stretch those cards’ lifespans.

But our experience also shows that MFG on Ada isn’t perfect. It’s still a mod in active development, and the unusual and reproducible input latency behavior we charted at 1440p on RTX 40-series cards is the sort of unexpected pitfall you might expect from an unofficial implementation of the feature. Cross your fingers that it’s an issue that can be fixed by the community.

The fact that MFG works as well as it does on Ada, even in this modded form, also makes us wonder whether Nvidia might just enable official support for it at some point, given the apparently bleak prospects for gaming graphics card pricing and future hardware generations as the AI boom shows no signs of abating. And such a move would provide much broader and more immediate performance relief for gamers than re-introducing ancient silicon like the RTX 3060.

Given that Nvidia is already working on bringing DLSS 5 to RTX 40-series GPUs, maybe this mod will convince it to throw in official MFG support, too. Fingers crossed.

Benchmarking Qwen 3.8 27B on RTX 5090 and beyond — VRAM capacity alone can't overcome severe software and inference engine bottlenecks

2026年9月8日 21:30

Alibaba’s Qwen 3.8 27B open-weight AI model came out a couple of weeks ago, and it immediately created a wave of hype among local AI enthusiasts thanks to its impressive intelligence benchmark results for a model of its size and capabilities.

Totaling around 17GB for four-bit quantized weights and offering built-in multimodal capabilities on top of its general aptitude, Qwen 3.8 27B immediately grabbed the attention of everybody with an RTX 5090, RTX 4090, or RTX 3090 (as well as a Radeon RX 7900 XTX, Radeon AI Pro R9700, or Arc Pro B70).

Were we on the verge of frontier-level intelligence from a four-bit quant on a single graphics card? Could everybody with a capable enough local AI setup go and cancel their Claude or ChatGPT subscriptions?

The answer, of course, as with every open-weight AI model hype cycle, is more complicated than just eyeballing the size of the model weights and comparing it to your available VRAM pool. Does the card or system you're using to host the model have enough VRAM left over to provide useful amounts of space for the model's context once everything is running? Do your host system and LLM inference engine deliver acceptable time-to-first-token, as well as high throughput beyond just bench-racing from an empty context window?

It's one thing if you just want to chat with a model and see what happens; it's another entirely if you want to put it to work, especially as impatient agents take the limits of human perception out of the picture.

We wanted to see what hardware and software stack Qwen 3.8 27B really wants in order to deliver solid performance, so we ran it on systems ranging from a desktop PC with discrete GPUs to systems with unified memory architectures like the DGX Spark, Mac Studio, and Ryzen AI Halo.

Our discrete GPU AI testbed includes the following components:

Tom’s Hardware Local AI Testbed

CPU

Ryzen 7 9800X3D

Memory

64GB (4x16GB) DDR5-5200

Motherboard

Asus TUF Gaming X670E-Plus Wifi

SSD

Corsair MP600 Pro XT 4TB

Power supply

MSI MPG Ai1600TS

Operating system

Ubuntu 26.04 LTS

Where it was possible to do so, we tested performance with Qwen 3.8 27B’s built-in multi-token prediction capabilities both enabled and disabled. Not all of the model runners we tested were able to support MTP within the amount of VRAM available to us on some of our platforms. We note where MTP was and wasn’t possible in our analysis of each platform, as well as in our charts.

RTX 5090 performance

We started with the RTX 5090, whose 32GB of GDDR7 and 1.8 TB/s of memory bandwidth would seem to make it an absolute no-brainer for getting the best local inference performance with this dense model. (Mixture-of-experts models tend to be friendlier to performance on lower-end hardware like the DGX Spark and AMD's Strix Halo, as their limited numbers of active parameters mean less data movement during inference).

As a baseline, we followed our usual local AI benchmarking approach: grab the latest build of llama.cpp from GitHub, build it, grab an Unsloth quantization of the model from Hugging Face, and run it. But our testing quickly ran into a speed bump.

RTX 5090 Llama benchmarks
Tom's Hardware
RTX 5090 Llama benchmarks
Tom's Hardware
RTX 5090 Llama benchmarks
Tom's Hardware
RTX 5090 Llama benchmarks
Tom's Hardware

Although llama.cpp will happily allocate the full 262K context length with this model on an RTX 5090, its processing speeds at long contexts on this card are dire.

Time-to-first-token with a single 5090 stretches to roughly 30 minutes, suggesting that something is just broken here. And tokens-per-second throughput drops far, far below what you would expect for having one of the world's fastest graphics cards at your disposal. No matter how you slice it, llama.cpp is not the right model runner for this hardware right now.

Next, we tried vLLM, a production-grade inference engine that's more at home in the data center than it is on the desktop, although it can comfortably serve in both roles—at least if your host system is up to its requirements. Even with 64GB of main memory in our test rig, we had to allocate another 64GB of swap just to let vLLM load Qwen 3.8 27B successfully for the first time. A lightweight stack this is not.

The vLLM maintainers provide an NVFP4 quantization of Qwen 3.8 27B and deployment recipes for both one and two RTX 5090s. We just so happen to have two RTX 5090s in the TH labs, so we were able to try out both configurations.

RTX 5090 VLLM

(Image credit: Tom's Hardware)

RTX 5090 VLLM

(Image credit: Tom's Hardware)

Serving Qwen 3.8 27B on one 5090 with vLLM certainly works in a pinch, but it's not ideal for long-context inference because the base recipe for it limits you to just a 32K context. To get the full 262K context, you really want a single card with more memory (like an RTX Pro Blackwell card with 48 or 72GB of RAM) or two 5090s, as we were able to test.

And a single card doesn't have enough memory to enable Qwen 3.8 27B’s built-in multi-token prediction (MTP), which is super helpful in getting faster decode performance from this setup. 20 tokens per second across the board without MTP is not an impressive baseline for a card of this caliber.

RTX 5090 VLLM
Tom's Hardware
RTX 5090 VLLM
Tom's Hardware

Get two 5090s into the picture, though, and decode speeds rocket upwards for vLLM (albeit at a high cost to prefill). 70-80 tokens per second across the context depth sweep is a fantastic result for a local setup, and TTFT remains fairly reasonable. But we can go faster.

Enabling MTP with vLLM gets us to 100-110 tokens per second on the decode side for only a small hit to prompt processing speed. This setup provides consistent performance at prompt processing speeds that don’t make you question whether something has gone seriously wrong. But it ought to be fast, because our dual RTX 5090 platform as tested here would currently ring in at over $13,000.

We also tried the SGLang inference engine on the RTX 5090 across similar configurations as we did with vLLM.

RTX 5090 Qwen 3.8 SGLang
Tom's Hardware
RTX 5090 Qwen 3.8 SGLang
Tom's Hardware

SGLang is much faster on a single 5090 for some reason – almost 3x faster than vLLM’s single-5090 recipe – and also ekes out a bit more context (37,740) versus vLLM. But if you want to get the full 262K that the model natively supports, you still need a second card or a different one with more VRAM.

RTX 5090 Qwen 3.8 SGLang
Tom's Hardware
RTX 5090 Qwen 3.8 SGLang
Tom's Hardware

Like vLLM, SGLang supports tensor parallelism across multiple GPUs, so enabling dual-GPU inference is as simple as adding another launch flag. And as with vLLM, there are a number of speculative decoding strategies you can add to the recipe to enhance output performance.

The takeaway from this first phase of testing: if you have a single RTX 5090 and don't need long-context inference from it, you can certainly get usable performance from one with this dense model. But you need to choose your model runner carefully.

And if you want the full context window, reasonable prompt processing times, and high throughput from Qwen 3.8 27B all at once, you really want a graphics card with more than 32GB of VRAM as a starting point (or multiples).

RTX 3090 and RTX 4090 performance

With the RTX 5090’s behavior settled, we turned to some older consumer cards to see how they handle Qwen 3.8 27B. The 24GB RTX 4090 and 3090 are evergreen favorites among local LLM fans thanks to their relatively large VRAM pools and relatively affordable prices on the used market, but as we've already emphasized, just being able to load the model weights is far from the whole picture.

These cards can fit the Q4_K_M GGUF of Qwen 3.8 27B with llama.cpp just fine, but they require using the Q8_0 quantization of the KV cache to fit the results in their smaller VRAM pools from the get-go, and they also require limiting the context depth to well under the model’s 262K native limit. We found that a context length of about 112K tokens was about the most we could get away with before running out of VRAM.

And unlike the 32GB RTX 5090, which can usually get away with having the Linux desktop window manager running next to the LLM and its infrastructure, these GPUs need every last byte of VRAM for the AI workload and nothing else. So you really want a separate graphics card at hand for these two cards if you're not running a headless server, which can introduce some setup headaches of its own as you discover how your particular motherboard handles PCIe slot bifurcation and enumeration of the primary graphics device.

RTX 3090 Qwen benchmarks
Tom's Hardware
RTX 3090 Qwen benchmarks
Tom's Hardware
RTX 3090 Qwen benchmarks
Tom's Hardware
RTX 3090 Qwen benchmarks
Tom's Hardware
RTX 4090 Qwen Benchmarks
Tom's Hardware
RTX 4090 Qwen Benchmarks
Tom's Hardware
RTX 4090 Qwen Benchmarks
Tom's Hardware
RTX 4090 Qwen Benchmarks
Tom's Hardware

Once you overcome those obstacles and get Qwen 3.8 27B up and running on these cards, llama.cpp exhibits the same performance cliff at long contexts on the RTX 4090 that we saw with the RTX 5090. But the RTX 3090 is oddly not affected. This suggests a bug somewhere.

We didn’t have time to dig into SGLang or vLLM behavior on these products, but given that you’re already tight for context on an RTX 5090, we’re doubtful that either of those inference engines would be an awesome way to run the model on these 24GB cards, unless you’re somehow ready to roll with multiple 3090s or 4090s from past acquisitions.

DGX Spark performance

Hardcore local LLM enthusiasts will scoff at the DGX Spark’s mere 27 GB/s of memory bandwidth for a dense model like Qwen 3.8 27B, and indeed, we've found that this platform isn't the fastest with dense models in our past testing.

But now that models like Qwen 3.8 27B support MTP with nothing more than a server launch flag, you can often get a major free boost to the decode speeds of platforms with limited memory bandwidth.

DGX Spark Qwen Benchmarks
Tom's Hardware
DGX Spark Qwen Benchmarks
Tom's Hardware
DGX Spark Qwen Benchmarks
Tom's Hardware
DGX Spark Qwen Benchmarks
Tom's Hardware

In our actual tests, the Spark's solid prefill processing performance means that it will often end up finishing inference turns at longer context lengths well before the RTX 5090 does with llama.cpp.

And beyond llama.cpp, the Spark is also well supported by SGLang and vLLM, so you can take advantage of those inference engines if they’re more to your taste. Consider also that a single Spark is still available for about $5000, and it’s a turnkey system that can be expanded into a handy cluster down the line if you want. So it shouldn’t be ruled out, even for serving this dense model.

Apple Mac Studio with M4 Max performance

The M4 Max-powered Mac Studio in our labs has the most memory bandwidth of any of the unified memory systems we have available, but as we've described in previous testing, that's only one metric that matters for local AI inference.

M4 Max Qwen Benchmarks
Tom's Hardware
M4 Max Qwen Benchmarks
Tom's Hardware
M4 Max Qwen Benchmarks
Tom's Hardware
M4 Max Qwen Benchmarks
Tom's Hardware

Prompt processing on this platform is slower than on Spark, so even if the Mac Studio can turn out more tokens than GB10 in the decode phase, it still ends up spending more time per inference turn than Nvidia's platform at longer contexts because that’s where it has to spend most of its processing time.

And at least in llama.cpp, using MTP on the Mac Studio actually causes a performance loss at shorter contexts for decode in exchange for a small boost at longer contexts, where it generally leads to improvements for other platforms. This demonstrates the value of actual benchmarking rather than spec-racing.

Ryzen AI Halo (Strix Halo) performance

AMD’s Ryzen AI Halo presents the worst-case performance scenario for this dense model: relatively low memory bandwidth and low prompt-processing performance.

Ryzen AI halo Qwen Benchmarks
Tom's Hardware
Ryzen AI halo Qwen Benchmarks
Tom's Hardware
Ryzen AI halo Qwen Benchmarks
Tom's Hardware
Ryzen AI halo Qwen Benchmarks
Tom's Hardware

Although MTP wakes up tokens-per-second throughput a bit on this system with llama.cpp, it can’t make up for the lengthy prompt processing times required for longer contexts. You can certainly run this model if a Strix Halo is the only box you have, but we’d seek out something more capable if you’re trying to do interactive long-context work.

Bottom line

When I first set out to explore Qwen 3.8 27B's performance, I figured this would be a relatively straightforward series of tests: plug in a single graphics card, load the model, get tokens, done. In practice, our experience required a lot more tinkering. And systems we might have initially written off as being not up to the task of running a dense model like this proved surprisingly useful.

In general, breathless claims of hundreds of tokens per second of throughput from an empty context window do not account for the full range of behavior one might see from an LLM on a given inference setup.

For just one example, whether it's down to a problem with (or just the expected behavior of) llama.cpp or something else about our software stack, the notion that you'd want to wait as much as 30 minutes or more for a response from Qwen 3.8 27B at long context lengths on an RTX 5090 is outrageous. But if you naively load Qwen 3.8 27B using llama.cpp right now, this is the experience you'll get.

Changing up inference engines is a natural next step, but there are trade-offs with that approach, too. You can load Qwen 3.8 27B on one 5090 using vLLM or SGLang, but those inference engines are much more conservative about the amount of usable context they’ll give you. The recipes we used only resulted in a context window of 32K tokens on a single 5090.

To enable the full 262K context length, we had to grab another RTX 5090 from the TH testing arsenal, at which point we got both great throughput and a TTFT sweep that could be considered interactive all the way out to the maximum context length from both model runners. But the price of replicating such a setup would exceed $13K right now.

You also might expect that a DGX Spark and its 273 GB/s of memory bandwidth wouldn't be useful for this dense model, but the prefill speed of the Spark ends up being fast enough that the TTFT remains relatively interactive even with a decode throughput of just 20 or so tokens per second with MTP, and that behavior holds out to the model's full native context length.

The M4 Max-powered Mac Studio has plenty of memory bandwidth on tap for decode, but its prompt processing speed means that the total time of an inference turn is dominated by that activity on this older Apple Silicon chip. The newer M5 Max and brand-new M5 Ultra would doubtless perform better, but we didn’t have those chips handy for this testing. And AMD’s Ryzen AI Halo gets the worst of it, with both low prompt processing speeds and relatively low TPS due to its memory bandwidth.

For all this, we really need to take a step back and consider the economics of local AI once again. $5K, $10K, or $15K or more for local AI hardware is a lot of tokens from leading-edge models at Anthropic or OpenAI (and even more from providers serving the recent slate of Chinese open-source models). A lot. And if time is money for you, barring compute constraints, those tokens will get back to you or your agent faster than anything you can run at home short of a DGX Station with its GB300 GPU.

So unless you’re working with sensitive data that requires on-premises processing, you’re an enthusiast who just wants to tinker, or you’re worried about the fate of open model distribution and inference more generally for some reason, you probably don’t need to rush out and build a box just for this model.

But if you do, be aware that delivered performance is more than just VRAM capacity or memory bandwidth, and that you might not get the best performance from your setup with the most common model runners like llama.cpp. Let experimentation and careful benchmarking lead you to the best results for your specific config.

We tested DLSS 5 in NBA 2K27 with every RTX 50-series GPU — first official release comes with a big performance hit, but almost every Blackwell card can run it at 1080p

2026年9月5日 19:00

Nvidia's DLSS 5 has arrived in NBA 2K27, and we've been up since midnight testing it across every RTX 50-series graphics card to see what the first official implementation of this tech can do and how it compares to the community-implemented mods that have been circulating over the past couple weeks.

Before we discuss the performance cost of DLSS 5, many will ask whether this tech is even worth getting excited about, given the intense controversy that it's sparked ever since Nvidia revealed it earlier this year.

In short: yes, absolutely.

This first-party implementation of DLSS 5, tuned under the full control of 2K Games' art directors and artists, looks incredible, full stop. If you see it running, you will want to leave it on. It doesn't look like "slop" or a cheap filter. It just looks correct, or at least more correct.

NBA 2K27 DLSS 5
2K
NBA 2K27 DLSS 5
2K

For just a couple of examples, without DLSS 5, even at ultra settings and with RT, the game has a slightly "plastic" or "flat" look. Faces can have a waxy uniformity that immediately indicates that you're looking at a video game, not a live broadcast. Hair looks totally and unnaturally flat. The whites of eyes can be unnaturally bright, giving a sort of googly-eye or doll-like effect that immediately broke my sense of immersion. I didn't think any of this would be a big deal for a sports game, but it all stands out.

NBA 2K27 DLSS 5 comparisons
2K
NBA 2K27 DLSS 5 comparisons
2K

DLSS 5 breathes incredible life into NBA 2K27. Skin looks like real flesh and blood. Eyes are rendered with the proper tones, depth, and sparkle. Hair looks a billion times more realistic. Every human on screen just looks more alive. You won't entirely forget that you're looking at a game, but it becomes much easier to suspend disbelief and get lost in the action.

My experience slapping DLSS 5 mods into games has produced promising but sometimes mixed results. If the polished experience that NBA 2K25 delivers is any indication, I can't wait for more studios to get their hands on it and integrate it into their games under the care of the same artists that created them.

Our testing methods

We did our best to deliver clean test numbers in the short time we've had with NBA 2K27. Performance in this can be tricky to measure. Benchmarking in the hub worlds or cutscenes will deliver lower performance than on the court, and even then, in-game events like fouls and time-outs will also result in skewed numbers if you're not careful. We picked a repeatable match-up from the game's career mode and sampled 60 seconds of continuous gameplay with DLSS 5 on and off, making sure to restart if we ran into any of the scripted events above.

We used NBA 2K27's maximum graphics settings as our baseline, including RT. We tested at native resolutions without upscaling, and we didn't test with Multi Frame Generation enabled. We view MFG as a cherry on top of an already solid baseline experience, not a baseline in itself.

Our test system is built with the following components:

Tom's Hardware 2026 GPU Test System

CPU

AMD Ryzen 7 9800X3D

CPU Cooler

Thermalright Phantom Spirit 120SE

Memory

32GB (2x16GB) G.Skill Trident Z5 Neo DDR5-6000 CL30

Motherboard

Asus TUF Gaming X670E-Plus Wifi

Storage

Inland Performance Plus 4TB PCIe 4.0 NVMe SSD

Power supply

MSI MPG Ai1600TS 1600W

Operating system

Windows 11 Pro

Graphics driver version

GeForce Game Ready 616.64

All of our performance results are captured using Nvidia's FrameView 2.0 utility. We measure graphics card power consumption directly with Nvidia's PCAT hardware power logging tool.

If you have any questions about our testing methods, let us know in the comments and we'll do our best to answer them. On to the numbers.

NBA 2K27 1080p performance

DLSS 5 performance

(Image credit: Future)

At 1080p, everything from the RTX 5070 on up is CPU-bound without DLSS 5 enabled, even with maxed-out settings and RT on. Turn on DLSS 5, though, and the differences in Tensor Core compute across the cards quickly becomes obvious. But even the RTX 5060 can manage nearly 60 FPS on average with DLSS 5 enabled at 1080p, so almost anybody with a Blackwell card can at least try out the feature.

Given these results, you might think to enable DLSS Super Resolution (aka upscaling). But because the amount of time needed to run the DLSS 5 model has a fixed cost per output frame that scales with your target resolution and largely dominates the total frame time, especially on lower-end hardware, you may find that enabling DLSS SR doesn't have as much of an effect on performance as we've come to expect. The game still has to wait for that final generative step to occur, even if DLSS SR cuts down some of the total frame time.

As noted, we didn't use MFG for these tests, simply because there's no need for us to make the numbers on these charts artifically large. If you do want to enable it, the multiplier you want will be determined by your monitor's refresh rate and your own personal tastes.

DLSS 5 performance

(Image credit: Future)

We did chart PC latency as estimated by FrameView, so you can get a sense of whether there's enough of a latency budget to enable MFG at all. At 1080p, every card is a solid candidate for enabling MFG without adding unreasonable amounts of input latency, although the RTX 5050 is in a borderline position. And all three of the 8GB cards might have trouble fitting the MFG model into their VRAM with these maxed-out settings and DLSS 5 enabled, so you might need to turn off RT at a minimum to use MFG on these lower-end GPUs.

DLSS 5 performance

(Image credit: Future)

As we already saw in early testing of the leaked DLSS 5 builds circulating in the community, enabling the model has a large impact on power consumption, likely due to the intense Tensor Core load required to run the DLSS 5 model.

NBA 2K27 1440p performance

DLSS 5 performance

(Image credit: Future)

Moving up to 1440p separates our cards some more. The RTX 5080 and RTX 5090 are still CPU-limited without DLSS 5. The RTX 5060 Ti 8GB, RTX 5060, and RTX 5050 all get a "Low VRAM" warning with these settings, so if you're trying to push this higher resolution with those cards, you might want to start tuning your DLSS upscaling settings to relieve VRAM pressure, even if it doesn't result in additional performance with DLSS 5.

DLSS 5 performance

(Image credit: Future)

But you really want an RTX 5060 Ti at a minimum here to get acceptable input latency with DLSS 5 enabled, and the RTX 5070 is the true baseline for a good experience. At the higher end, the RTX 5070 Ti and RTX 5080 both provide fluid frame rates even without MFG, and you have the latency budget to enable framegen without worry on anything from the RTX 5070 on up.

DLSS 5 performance

(Image credit: Future)

DLSS 5's power consumption on all these cards (except the chugging RTX 5050) rises as expected with the higher-resolution output frame we're asking it to generate. Everything from the RTX 5060 up to the RTX 5070 Ti ends up running at its power limit, while the RTX 5080 and RTX 5090 still have room to stretch out. The unconstrained MSI RTX 5090 Lighting Z sucks down a ton of power to deliver its scorching performance.

NBA 2K27 4K performance

DLSS 5 performance

(Image credit: Future)

Running NBA 2K27 at 4K with DLSS 5 is extremely demanding, since the model has to generate an output frame with more than twice as many pixels than at 1440p. Even the RTX 5070 Ti is straining here, as its input latency with DLSS 5 is potentially too high to enable MFG while still delivering a responsive gameplay experience.

DLSS 5 performance

(Image credit: Future)

You can see why Nvidia only recommends the RTX 5080 and RTX 5090 for a 4K DLSS 5 experience in this title, as they're the only two cards that deliver high enough baseline performance with low enough input latencies to make MFG practical. And even then, the RTX 5080 is on the edge of what we'd consider an acceptable input latency before enabling MFG.

DLSS 5 performance

(Image credit: Future)

Our power consumption chart at 4K is a bit of a mess at the low end since the RTX 5050 and RTX 5060 are crushed by the demands of DLSS 5 and end up chugging. Really, though, we're here to see how the RTX 5080 and RTX 5090s handle this incredibly Tensor Core-intensive work. Both cards end up running at their power limits.

The MSI RTX 5090 Lightning Z shows how much power scaling the GB202 GPU has left in it once you remove the constraint of a single 12V-2x6 connector. We were wondering whether the power numbers we saw from a modded version of Control were a fluke, but DLSS 5 in NBA 2K27 goes even harder and pushes the RTX 5090 Lightning Z to nearly 850W on average in this test. For 48% more power than the RTX 5090 Founders Edition, you get 22% higher performance.

On a big 4K screen like an OLED TV, a fluid 90 FPS at a native 4K resolution with the level of detail and realism that DLSS 5 adds is an astounding gaming experience. You really haven't seen anything like it. But mere mortals with single-plug 5090s will probably want to enable MFG 2X at a minimum.

Bottom line

It's early days for DLSS 5 performance, and in its first official showing in NBA 2K27, it certainly has a large performance cost—sometimes well over 50% on lower-end hardware. But the generational leap in realism it provides for every human on screen is well worth it to my eye, at least. Now that I've seen a first-party, developer-driven implementation of DLSS 5, I want it in every game where its photorealistic enhancements would make sense, without question.

And as it's implemented in NBA 2K27, our performance results show that the tech is accessible enough that almost anybody with an RTX 50-series graphics card can try it out and still enjoy a fluid and responsive experience, even without Multi Frame Generation. Most of us aren't playing on 4K monitors, and at 1080p and 1440p, the Blackwell card you may already have is probably up to the task of running DLSS 5 with acceptable baseline performance.

DLSS 5 does upend some intuitions we've developed around the interactions of upscaling and frame rate. The large fixed frame-time cost that's required to generate the DLSS 5 output frame is entirely dependent on your target output resolution, making the lowered input resolution of DLSS 4.5 upscaling far less of a performance multiplier than it might otherwise be.

Given those realities, DLSS MFG can still make for a more fluid experience alongside DLSS 5 with a couple of clicks in a menu—at least assuming you have enough VRAM to hold the game assets, the DLSS 5 model, and the MFG model all at once. On 8GB graphics cards, this may require some tweaking to get working.

Nvidia has promised to continue improving the fundamental performance of the DLSS 5 model with time, and the fact that it's moved from requiring a dedicated RTX 5090 to run in the demos we saw at GTC to running on an RTX 5060 today is a sign that such promises do bear fruit.

Boosting the fundamental performance of the DLSS 5 model is vital, because as we all well know, prices for most Blackwell cards have spiked across the board to the point that former midrange options like the RTX 5060 Ti 16GB and RTX 5070 are many hundreds of dollars more expensive than their MSRPs and well out of whack with their former performance-per-dollar propositions. A hardware upgrade to get better performance is no longer a no-brainer for many PC builders, and that problem is going to get even worse as the AI boom continues unabated.

Those elevated prices are also why Nvidia's post-launch pledge to bring the DLSS 5 model to RTX 40-series cards later this year is a welcome development. In our limited experience with community mods, Ada GPUs run the current DLSS 5 model about as well as Blackwell cards do, and gamers who haven't already upgraded from their 40-series cards certainly won't be raring to move off that hardware any time soon. Locking DLSS 5 to new GPUs that are prohibitively expensive to buy for non-technical reasons isn't a winning strategy for goodwill or broad adoption, and it's good to see Nvidia acknowledge that reality.

For all that, we shouldn't lose sight of the fact that DLSS 5 is a huge, exciting leap forward in the never-ending pursuit of photorealism in real-time graphics. When you toggle it on and off in NBA 2K27, it feels like a generational improvement in rendering technology at the press of a button. Now that I've seen it in action, I don't want to play NBA 2K27 without it, and I hope to see more first-party integrations of it soon.

Nvidia PAIR utility joins every GPU in your home into a cluster for agentic AI tasks — tool uses spare cycles to keep agent swarms from hammering one GPU

2026年9月4日 00:00

If you're a token-hungry AI enthusiast, and if you or your family happen to have PCs with idle GPU cycles to spare in this economy, Nvidia wants to make it possible to harness those cycles so you can save cash on cloud tokens and keep your work private. At IFA 2026, the company is introducing a local distributed AI clustering tool called the Personal AI Router (PAIR) that dispatches agentic AI sub-tasks from your main PC to systems on your home network that have suitable GPU cycles to spare.

As Nvidia tells it, when a user runs a local AI agent and gives it a goal to complete, that central agent might then spawn several sub-tasks carved out of that larger goal. If those sub-tasks or sub-agents are all running on the same GPU, the contention they create might cause the task to finish more slowly than it could if each sub-agent had a dedicated compute node to work with.

PAIR is a tool that can make that distributed AI work happen on a home network. Assuming that a family or shared household is sufficiently flush with idle GPU resources, PAIR canYEa assign each participating system one of those sub-tasks to perform and return the results to the main node, potentially resulting in faster completion of the larger agentic task.

Of course, systems on your local network won't always be idle. Their owners will frequently use the GPUs in their systems for gaming, creative work, or AI tasks of their own. If a user needs their GPU back, PAIR purports to gracefully deal with those changing conditions. It doesn't reserve dedicated capacity from other PCs; it's elastic by design and will make the best of the resources available to it at any given moment.

This unpredictable availability of spare cycles does, of course, mean that quality of service is not assured from a PAIR cluster. But for long-running tasks that don't need to be done on a strict deadline, being able to put spare compute to work could still be more effective than running an agent swarm on a single node.

PAIR sounds relatively simple to set up. It creates a proxy for popular AI front-ends like LM Studio and Ollama to connect to. PAIR then orchestrates work across available nodes on the network and returns the results of that work to the originating application on the head node.

In turn, participating PAIR nodes also need to be running Ollama or LM Studio and have a PAIR installation of their own. Nvidia says that enrolling systems in a PAIR cluster is straightforward and relies on mDNS or an IP address fallback for discovery. PAIR will also help initiate model downloads on participating systems, but Nvidia says that nodes don’t need to have identical models or sets of models downloaded to participate.

If more systems do have a given model available, though, it broadens the pool of potential nodes that can handle a request if the orchestrator agent needs a particular model’s capabilities.

PAIR will run on any DGX Spark (or other GB10) box, as well as GeForce RTX 20-series graphics cards or newer. It also supports Macs with M4-series processors or newer for inference. Accordingly, the PAIR client will be available for Windows, macOS, and Linux.

We explored early DLSS 5 performance with community mods — and the limits of the 12V-2x6 power connector may hold it back on the RTX 5090

2026年9月2日 21:21

It's been a wild few days for gamers, as a version of the DLSS 5 model leaked with NBA2K27 and promptly got modded into every game under the sun. After initial builds that relied on ReShade to make the model work, the community has since wrapped up DLSS 5 into an Optiscaler package that makes adding it to most games nearly painless. With that development, we wanted to get a sense of how the tech performs ahead of its official release.

We're not going to weigh in on the aesthetic or philosophical implications of applying DLSS 5 to a particular title here. That ground has been extremely well trodden already, and if you haven't been stuck under a rock this past week or so, you've likely seen what DLSS 5 can do for yourself.

We're more interested in getting an idea of the performance cost of getting DLSS 5 running, along with the power requirements it incurs. Whether you love or hate this model's effect on a game's appearance, you can weigh your feelings against the performance cost involved in getting there.

Before we get to the numbers, some caveats: the Optiscaler DLSS 5 injection method may not represent the behavior of this model when developers integrate it through Nvidia's Streamline framework. Streamline may offer less overhead and better performance than what we saw here.

But Nvidia's technical paper on DLSS 5 says that the model takes 8 ms to execute on a 4K frame, and that's essentially identical to the performance metrics that we saw from Optiscaler when it's running. So it seems unlikely that the official version will be far off the performance we measured if it's implemented in these titles using the proper channels.

For these limited tests, we chose to focus on the RTX 5090 Founders Edition's behavior under DLSS 5 to explore the largest performance drop one might expect from this community implementation of the tech.

Out of curiosity, we also grabbed MSI's RTX 5090 Lightning Z from the TH arsenal. This RTX 5090 has a unique two-12V-2x6 connector power setup and a 1000W power limit available through its Extreme vBIOS. Given the apparent arithmetic intensity of DLSS 5, we wanted to see whether running it would consume any of the extra available power headroom from that card's twin-plug design.

For our small sample of games, we used the following basic settings: a 4K output resolution target fed by DLSS Performance upscaling (i.e., a 1920x1080 input) without frame generation, as well as maximum raster and ray-tracing quality settings (or path tracing, in the case of Cyberpunk 2077) and DLSS Ray Reconstruction where it was available.

We didn't use Multi Frame Generation in these tests. We want to cleanly represent the performance baseline you can expect from DLSS 5 given the above constraints.

Cyberpunk 2077 performance

We started our performance explorations with Cyberpunk 2077. Optiscaler was an easy install for this game, and we confirmed that neural rendering was making the expected difference in facial detail and lighting realism when we activated it.

Cyberpunk 2077 DLSS 5 performance

(Image credit: Future)

In exchange for those enhancements, the RTX 5090 Founders Edition takes a 42% hit from DLSS 5. Although the Lightning Z takes a slightly smaller 39% hit, its baseline performance is higher than the Founders Edition, so the MSI card maintains a 60 FPS average and much higher 1% lows than the single-connector card. That difference in smoothness is evident.

Cyberpunk 2077 DLSS 5 performance

(Image credit: Future)

Part of that difference in performance comes down to power. As we noted, the RTX 5090 Lightning Z has a much larger 1000W power limit to play with, and it puts that headroom to use here.

The jump in power consumption from DLSS 5 on the Lightning Z is eye-popping: 580 W to 723 W, or a 25% increase. The RTX 5090 Founders Edition only goes from 482 W to 551 W, which means that it's practically power-limited.

Hogwarts Legacy performance

Unlike Cyberpunk 2077, Hogwarts Legacy doesn't implement path tracing, but its extensive RT effects are still impressive. Importantly for our purposes, it's easy to get Optiscaler into this title.

Hogwarts Legacy DLSS 5 performance

(Image credit: Future)

Hogwarts exhibits the largest performance drop of these three games on the RTX 5090 FE, at 49%. The Lightning Z card drops just 42%, though, and its higher baseline performance makes for a smoother experience.

Hogwarts Legacy DLSS 5 power

(Image credit: Future)

The Lightning Z once again illustrates just how much power DLSS 5 can pull if it's available, even in this "lighter" RT title. Without DLSS 5, the MSI card draws 480 W on average. With neural rendering enabled, it pulls 720 W, or a whopping 50% more power.

Control performance

Finally, we tried the game that started the community DLSS 5 wave last week: 2019's Control. This early RT title still poses a stiff challenge for modern hardware if you crank its RT settings to the absolute max, as we did here.

DLSS 5 Control performance

(Image credit: Future)

The RTX 5090 Founders Edition takes another 42% hit to performance under these conditions, and the Lightning Z also drops 42%. But the higher power limits of the MSI card let it deliver 20% higher average frame rates and much higher 1% lows than the Founders Edition, and that's a smoothness boost you can really feel.

DLSS 5 Control power measurements

(Image credit: Future)

That extra performance does come at the expense of significantly higher power draw. The Lightning Z is already pulling a shocking 691 W without DLSS 5, but enable neural rendering, and it jumps to 802 W. The Founders Edition just runs into its 575 W power limit again.

Bottom Line

At least in our experience with the community-made mod packages available so far, DLSS 5 has a large performance cost, similar to what we saw with ray tracing before the wide availability of DLSS: in the range of 39% to 49%, according to our small sample of results so far.

Official versions of DLSS 5 may perform better, but we won't know for sure until later this week. In any event, our observations of the behavior of the tech as implemented by the community match up with Nvidia's published guidance so far.

We don't expect vastly different or better performance from first-party implementations, although a developer's official DLSS 5 recipe could certainly look better in a given title than the current free-for-all with sliders that were never meant to be exposed to end users. We'll have to see what happens as developers take the time to tune DLSS 5 for their games as the tech makes its way into more titles through official channels.

Officially, DLSS 5 only supports RTX 50-series graphics cards, and if you're interested in running it with RT or PT, our results show that you're likely going to need to enable the full stack of Nvidia software tech available from Blackwell: DLSS Performance or Ultra Performance to achieve solid baseline frame rates, plus Multi Frame Generation for acceptable output fluidity.

Some have suggested that you don't need to run RT or PT alongside DLSS 5, but that doesn't make any sense at all given our experience thus far. Applying DLSS 5 to purely rasterized input does make it look better than it otherwise would, but it makes the higher-quality input of RT or PT titles look even better still. You definitely don't want to give up those advanced rendering techniques just because DLSS 5 exists.

But those software factors aren't the biggest hurdle that might affect DLSS 5's delivered performance, at least at the extremes. We were surprised to find that the RTX 5090 Founders Edition design is probably holding back the performance of DLSS 5 on that card, even with its 575W TGP. The Founders Edition was consistently at or near its power limit in our tests with neural rendering enabled, which tracks with the demonstrated evidence that, like other image generation models, DLSS 5 is extremely demanding of the Tensor Cores that power its underlying diffusion model.

The exotic MSI RTX 5090 Lightning Z and its 1000W power limit show that when it's stacked on top of ray tracing or path tracing, DLSS 5 can and will happily slurp down hundreds more watts than the Founders Edition—or any other RTX 5090 with a single power connector—is built to provide.

And when you pair that fact with its large performance cost on today's hardware, it's clear that DLSS 5 is a multi-generational technology built to take full advantage of graphics cards that don't exist yet, whether they're RTX 50 Super-series products with higher power limits or future GeForce products built with more advanced Tensor Cores on a newer silicon process node.

Even as it stands, the results we've seen from DLSS 5 so far in both official demos and in its rapid proliferation through community mods suggest that neural rendering techniques like this will mark a new epoch of photorealistic fidelity for real-time graphics and new generations of hardware that are better suited to running demanding AI models like this in real time.

It's an exciting new frontier out there, even if things feel a bit like the Wild West for DLSS 5 right now. Hold on to your hats.

❌
❌