Slow wired upload speed vs Linux on same hardware

Dawid Oosthuizen 26 Reputation points
2020-09-09T05:18:31.053+00:00

Intel® Ethernet Controller X550-AT2, 10G network interface on ASRock Rack ROMED8-2T with AMD EPYC 7232P processor.
Windows 10 Pro for Workstations, 2004.
Latest Windows updates and Intel drivers installed as of 09/09/2020.

This machine is a dual boot, the above Windows version, and Ubuntu 20.04.
When doing a speed test, I get good performance from Ubuntu, but very poor uploads from Windows. This is on the same machine, the exact same hardware.

The WAN link is 1000Mbps down, 50Mbps up.

This is the Windows speedtest result:
23320-windows-10-speedtest.png

This is the Ubuntu speedtest result:
23325-ubuntu-2004-speedtest.png

I have tried to tweak the adapter's advanced driver settings in Windows, such as disabling LSO, etc. No luck, performance remains poor.

I've also noticed it on another PC running Windows 10 Pro, and a laptop running Windows 10 Pro for Workstations, both give the same poor upload performance. Whereas my other Ubuntu 20.04 Server machine, and also my phone connected via Wi-Fi, is getting good upload speeds.

I have even taken the Windows laptop and plugged it straight into my incoming WAN connection (bypassing router), and it still gets poor upload speeds.

Incidentally, when the speed test is running, I can see that the upload looks bursty on Windows, like it is only getting chunks of data here and there, while in Linux and on Android it looks the same as the download, the graph is drawn at a consistent high rate and with consistent high values.

Windows for business | Windows Client for IT Pros | Networking | Network connectivity and file sharing

50 answers

Sort by: Newest
  1. Gary Nebbett 6,536 Reputation points
    2020-12-20T21:06:25.78+00:00

    Hello @mensa84 et al.,

    I think that I understand at an abstract level what is going wrong. In the trace there are 127 retransmissions: 94 fast retransmits and 33 normal retransmits. Each retransmission causes the "congestion window" to close a little (and, therefore, the throughput to be reduced). Most of the "fast retransmissions" are "spurious" (were not actually needed). WireShark "Expert Information" shows:

    49852-image.png

    Both "fast retransmission" and "spurious retransmission" are mentioned. Here is a more detailed view of what is happening:

    49772-image.png

    At 22:18:14.967353 the byte range 172967:174387 is transmitted for the first time.

    Between 22:18:14.997584 and 22:18:14.997586 (30.231 milliseconds after the above and within a 2 microsecond period), 12 acknowledgements appear in the event trace (they have probably been "collected" by a lower level and then "indicated" as a group).

    At this point in the trace, one estimate of "Round Trip Time" (RackMinRtt) is 18.994 milliseconds. Either the three occurrences of 172967 as the next "unacknowledged" sequence number (simple duplicated acknowledgement threshold) or a combination of this with "recent acknowledgement" ("rack") timings suggest that a "fast retransmission" would be apposite. This happens before the next packet is examined a few nanoseconds later, showing that the "missing" byte range has been acknowledged.

    These images of the Microsoft-Windows-TCPIP trace data show that the packets arrive as a group and that, subsequently, a fast retransmission occurs:

    49853-image.png
    49735-image.png

    It is the high degree of "out-of-order" delivery of TCP segments (i.e. segments that are not received in the same order/sequence as which they were sent in) which perhaps limits the number of users who experience this behaviour.

    I don't yet fully understand what can be done to influence this behaviour. Setting (in the registry) TcpMaxDupAcks to 3 might help (see https://learn.microsoft.com/en-us/troubleshoot/windows-server/networking/description-tcp-features) but, beyond that; I don't currently have any suggestions.

    Another question is whether this behaviour has any impact on one's typical workload. It certainly substantially slows "streaming" of data to a server (as happens when posting data to a web server - this is what many "speed tests" measure), but it might have less (or no) impact when uploading data to a file share using SMB encapsulated in a VPN connection (in a "working from home" scenario) - the command/response nature of such transfers (and perhaps the fact that the "top level" protocol is Encapsulating Security Payload (ESP) rather than TCP) might slow down all devices/OSs to approximately the same level.

    Gary

    Was this answer helpful?


  2. Moosa Mahsoom 1 Reputation point
    2020-12-20T20:37:55.237+00:00

    I am receiving a very similar (actually looks identical) issue on my Windows 10 with both my Ethernet and WiFi connections on my PC. However, If i use linux (via bootable USB), the issue doesn't persist. the issue didn't exist before but around 2 weeks i started noticing it. It became a major issue when i was streaming to my youtube channel and my upload speed massively tanked. Usually I stream at 10k bitrate but this time it was struggling to send even 2k.

    Was this answer helpful?

    0 comments No comments

  3. Gary Nebbett 6,536 Reputation points
    2020-12-18T13:30:36.743+00:00

    Hello @mensa84 ,

    I need to do some more checking to verify whether this is always or mostly the case, but I think that there are hints of the cause of the problem in the trace data. Here is a short extract:

    49551-image.png

    At some time before this extract, the 1420 bytes between 172967:174387 were sent from your system. In the extract above, we can see that the bytes up to 172967 were acknowledged and, initially, some bytes after 174387 were acknowledged with a "selective acknowledgement" (sack) - a warning that the bytes between 172967:174387 may have been lost.

    A little later (at 14.997585, highlighted in blue) the bytes between 172967:174387 are finally acknowledged.

    Now, very oddly, 100 microseconds later (at 14.997685, highlighted in yellow), your system retransmits the bytes between 172967:174387. If the timings in the trace are accurate and inferences drawn from the trace data match the Windows TCP/IP implementation's understanding of the state of affairs, there is no reason to retransmit the data.

    The retransmission of the data is not, in itself, a major problem but the fact that a retransmission was "necessary" causes the "congestion window" to be reduced. The congestion window is dynamically updated throughout the transfer but never increases to anywhere near the size needed to get the full "nominal" throughput of the link.

    My suspicion is that there is some "receive offloading" taking place - if that is the case and is the cause of the slow performance, we might be able to disable it. I will do some more analysis to increase my confidence level that this is the problem.

    Gary

    Was this answer helpful?

    0 comments No comments

  4. mensa84 6 Reputation points
    2020-12-17T21:23:11.48+00:00

    Hello,

    thanks for the explanation!
    I did this now, here is the SlowUp.etl zipped: https://drive.google.com/file/d/1rt_jvuq-zr4XPu-OV4bth8iLYMJY08i_/view?usp=sharing

    Do you see anything?

    Was this answer helpful?


  5. Gary Nebbett 6,536 Reputation points
    2020-12-11T13:31:29.307+00:00

    Hello @mensa84 ,

    Open a PowerShell window running as administrator. This image shows one way of doing that (right mouse clicking over the Windows PowerShell icon and choosing "Run as Administrator"):

    47443-image.png

    Now get ready to run the speed test (but don't start it yet) and then paste the following four commands into the PowerShell window:

     New-NetEventSession -LocalFilePath $Env:TEMP\SlowUp.etl -Name SlowUp  
     Add-NetEventPacketCaptureProvider -TruncationLength 100 -Level 255 -SessionName SlowUp  
     Add-NetEventProvider -Name "Microsoft-Windows-TCPIP" -Level 255 -SessionName SlowUp  
     Start-NetEventSession -Name SlowUp  
    

    Now start the speed test. As soon as it ends, paste the following two commands into the PowerShell window:

     Stop-NetEventSession -Name SlowUp  
     Remove-NetEventSession  
    

    You should now have a file containing the trace data (%TEMP%\SlowUp.etl, i.e. a file named SlowUp.etl in your temporary files directory).

    The file will be quite large, and the best way of sharing it is by copying it to a network storage service, such as Microsoft OneDrive, Google Drive, Dropbox, etc., and posting a link to the file.

    Gary

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.