Repository navigation
Improve performance and reduce latency of Transport.write by attempting to send data immediately if all write buffers are empty - #619
Conversation
|
This also increases aiohttp websocket echo performance by 20-25% |
be094bc to
96dc94c
Compare
|
Hi @fantix, I saw you recently commited something to uvloop. Do you know if someone could take a look at this PR and maybe give a feedback? |
|
Yeah I'll go through all issues/PRs again and include some in the final 0.21 release, it's just all taking some time unfortunately, thanks for the PR and your kind patience! |
…is removes unnecessary syscall
Hi @fantix, sorry to bother, I know it is open source, voluntary work and you're very busy. But would you have some time to look at this (and other my PRs) any time soon? Also, 0.21.0 is already released :) |
|
Can someone from the contributors have a look at this PR? It would be nice to improve the performances! |
|
@MagicStack guys please finish this CR. |
|
Updates on this? @fantix |
|
@fantix are you planning to review and release this soon? 🙏 |
|
@1st1 Is there any news about this PR ? |
|
Sorry about the delay! I'm back on this one now. |
fantix
left a comment
There was a problem hiding this comment.
Actually instead of replicating the fast-path code, maybe we can just relax the fast-path condition, like:
@@ -424,7 +424,7 @@ cdef class UVStream(UVBaseTransport):
cdef inline _initiate_write(self):
if (not self._protocol_paused and
(<uv.uv_stream_t*>self._handle).write_queue_size == 0 and
- self._buffer_size > self._high_water):
+ (self._buffer_size > self._high_water or len(self._buffer) == 1)):
# Fast-path. If:
# - the protocol isn't yet paused,
# - there is no data in libuv buffers for this stream,(kind of a follow-up of 46dd8f3)
I believe this would have the same performance boost as this PR, yet being compatible with the current code base.
|
Hi @fantix, The last one gives a decent boost, I would definitely suggest to merge it Also slightly irrelevant here, but any chance we can figure out why tests are sporadically failing? For my PRs I was just re-running tests 2-3 times until they pass :) |
|
Roger, I will get to them soon. I'm spotting at least 2 more flaky tests, I shall fix them first (or skip non-critical ones for now). |
Done.
|
| if used_buf: | ||
| PyBuffer_Release(&py_buf) |
There was a problem hiding this comment.
I added PyBuffer_Release for consistency, but I think that the outer check if blen == 0: is not really necessary.
We already filter out empty data objects twice. First in write(), second in _buffer_write().
We can probably remove the check in write() and leave check only in _buffer_write(). Both write and writelines rely on it.
I can do it, but I'd rather put it into a separate PR
Changes ======= * Add Python 3.15 and 3.15t wheel builds and CI coverage (MagicStack#758) (by @honglei @fantix in f7c0547) * Add support for the `eager_start` keyword argument in create_task() (MagicStack#748) (by @samypr100 in 3cbb095 for MagicStack#746 MagicStack#718) * Add thread name prefix to the default thread pool executor (MagicStack#636) (by @inikolaev @fantix in 0582f94 for MagicStack#562) * Add support for special hostname `<broadcast>` (MagicStack#592) (by @jpbede in 3060ceb for MagicStack#540) * Upgrade libuv to v1.52.1 (MagicStack#753) (by @fantix in e8efea4 for MagicStack#752) * Improve performance by using Python C API to enter/exit context (MagicStack#627) (by @tarasko in 837ef22) * Improve performance/latency of Transport.write (MagicStack#619) (by @tarasko in 1d9b6e0) * Replace some SSL vectorcall with direct methods (MagicStack#626) (by @tarasko in 6a27cbe) * Optimize SSL buffered reads using C values (MagicStack#629) (by @tarasko in a308f75) Fixes ===== * Detach socket on create_connection cancellation to prevent fd double-close (MagicStack#740) (by @junjzhang @fantix in dc680eb for MagicStack#645 MagicStack#738) * Remove `loop._ready_len` in favor of `len(loop._ready)` (MagicStack#721) (by @x42005e1f in 5910a18 for MagicStack#720) * Prefer inspect.iscoroutinefunction (MagicStack#705) (by @MatthieuDartiailh in fd65027 for MagicStack#703) * Fix context tests by explicitly yielding after run_in_executor (MagicStack#743) (by @samypr100 in 6cd24cb) * Fix test_create_connection_open_con_addr with Python 3.13.9+ (MagicStack#713) (by @shadchin in b93141a for MagicStack#701) * Skip flaky test_cancel_post_init on asyncio 3.13+ (MagicStack#714) (by @fantix in 3ea5c85 for MagicStack#709) * Fix flaky test_fs_event (MagicStack#717) (by @fantix in 8da4547) * Use C __atomic builtins for debug counters (MagicStack#719) (by @fantix in 836e3b2) * Update the example to work with Python 3.14 (MagicStack#710) (by @Jamie-Chang in b74c2f1) Build ===== * Replace pkg_resources with packaging and use Cython 3.1 (MagicStack#742) (by @samypr100 in b377b7c for MagicStack#729) * Remove wheel as a build dependency (MagicStack#696) (by @DimitriPapadopoulos in 173e88c) * Support any number of flags in UVLOOP_OPT_CFLAGS (MagicStack#630) (by @mgorny in 963a5f3)
Current implementation of UVStream.write almost never sends data immediately. Instead the data is stored in the buffer and picked up later by uv_check callback.
This introduces unnecessary latency and CPU overhead.
Despite being a change only in UVStream it also directly benefit _SSLProtocolTransport.write latency since it uses UVStream.write to send ssl frames.
This PR increases RPS rate between echoclient and echoserver by roughly 10%.
echoclient --worker 1 --num 200000echoserver --uvloop --protoBecause of this change a couple of test had to be tweaked.
Please note that the current implementation of asyncio does the same thing. It tries to write directly to the socket, and, only if EWOULDBLOCK happens, the data is added to the buffer
https://github.com/python/cpython/blob/c13e7d98fb8581014a225b900b1b88ccbfc28097/Lib/asyncio/selector_events.py#L1065
Apart from that:
I have relaxed fast-path condition further and also removed self._buffer_size > self._high_water in _exec_write. Seems to be working fine.
Added primitive return types to _try_write and _exec_write. Previously there were PyObject_RichCompare calls for values returned by _try_write.
Simplied return value meaning for _try_write. Now it just returns number of bytes written or -1 in case of fatal error.
Removed some signed/unsigned and Py_ssize_t -> int conversions in _exec_write and try_write. Just use Py_ssize_t without conversion when possible