feat(posix): add posix_descriptor for adopting native descriptors - #372
Open
mvandeberg wants to merge 1 commit into
Open
mvandeberg wants to merge 1 commit into
mvandeberg wants to merge 1 commit into
Conversation
|
An automated preview of the documentation is available at https://372.corosio.prtest3.cppalliance.org/index.html If more commits are pushed to the pull request, the docs will rebuild at the same URL. 2026-09-28 19:59:33 UTC |
|
GCOVR code coverage report https://372.corosio.prtest3.cppalliance.org/gcovr/index.html Build time: 2026-09-28 20:15:35 UTC |
mvandeberg
force-pushed
the
pr/posix-descriptor
branch
2 times, most recently
from
September 28, 2026 19:08
92bcedf to
6a09190
Compare
posix_descriptor adopts an already-open pollable file descriptor — a
character device, inotify, eventfd, timerfd, pidfd, pipe, tty, or a socket
kind corosio does not otherwise wrap — and drives it from the io_context
with read_some, write_some and wait(). It never creates a descriptor; the
caller supplies one. It ships on epoll, kqueue, select and io_uring, with
a devirtualized native_posix_descriptor<Backend> twin.
The name is platform-qualified deliberately. Portability comes from the
interfaces, not the name: a posix_descriptor is an io_stream, so
capy::read/write, capy::Stream-constrained algorithms and TLS layering
work on it exactly as they do on a socket.
Two contracts shape the implementation:
- assign() validates before it mutates. A rejected descriptor leaves the
object holding whatever it held before, pending operations included,
and leaves ownership of the descriptor with the caller. Regular files,
block devices and directories are rejected with
operation_not_supported; stream_file and random_access_file adopt
those. The file-type test is a reject-list rather than an accept-list
because eventfd, timerfd, inotify and pidfd are anonymous inodes whose
st_mode type bits are all zero, so an accept-list would reject exactly
the kinds this type exists to carry.
- O_NONBLOCK is applied lazily at the first read_some/write_some, never
at assign(), and never restored. wait() never touches it, so adopting
STDIN_FILENO to await readiness cannot flip the parent shell's
terminal to nonblocking. The flag lives on the shared open file
description, which is also why restoring it would race every other
holder, and why a dup() is no escape.
On io_uring the lazy flag is a cancellability requirement rather than only
a contract: a blocking read punted to an io-wq worker cannot be cancelled.
READV/WRITEV are submitted at offset -1 so the kernel uses the
descriptor's own file position, and an EAGAIN completion arms a poll_add
and resubmits. That second submission opens a window in which a cancel can
land between the EAGAIN completion and its dispatch, invisible to both
cancel-by-fd and a one-shot stop_callback, so the implementation carries
two generation counters — one bumped by cancel(), one by any descriptor
change — snapshotted at prepare time and compared before the re-arm. They
are separate because the two abandonment reasons name different codes to
the caller.
Three defects in shared reactor code surface through non-socket
descriptors and are fixed here:
- epoll delivers EPOLLHUP alone for a pipe or tty whose peer closed, and
that bit mapped to no reactor event, so a parked read never ran and
the edge-triggered registration never repeated: io_context::run() hung
forever. EPOLLHUP now widens to read and write readiness. Sockets
already carry both bits alongside EPOLLHUP, measured.
- getsockopt(SO_ERROR) fails with ENOTSOCK on a non-socket, and that
errno reached the caller in place of the kernel's real error. On
ENOTSOCK the probe now yields, and the dispatch is forced to re-run
each parked operation's own syscall so EPIPE or EIO surfaces instead.
- io_uring treated POLLHUP in a poll's revents as a fault and
substituted EIO, while the reactors treat it as readiness. POLLHUP is
now readiness for wait(read) and wait(write); only POLLERR and
POLLNVAL name an error, and wait(error) is unchanged.
The devirtualized twin's read_some, write_some and wait awaitables
hand-rolled an await_ready/await_resume pair that predates op_base, testing
the stop token directly. That reported a cancellation on an operation the
backend had already completed successfully. All three now derive from
bytes_op_base and void_op_base like every other awaitable, so a stop request
arriving after completion no longer displaces the result.
stream_file::assign() and random_access_file::assign() now validate before
mutating, matching the descriptor contract, and cancel in-flight work
before closing the file they hold — the POSIX pool reads fd_ and offset_
at execution time, so an uncancelled operation would otherwise complete
against the newly adopted file and report success.
This breaks one shipped behavior: both file types now reject a pipe, which
they cannot position with preadv/pwritev, and posix_descriptor is the home
for those. IOCP is unchanged; its assign() still closes the held handle
before registering, and the docstrings now say so.
Adds a guide page with compiled snippets covering adoption, the dup() and
descriptor-flag rules, what is rejected, and the SIGPIPE limit — a write
to a descriptor whose peer has closed raises it, unlike the socket types,
because MSG_NOSIGNAL is a send() flag with no writev equivalent. Two
backend asymmetries are documented rather than papered over: where a
kernel refusal surfaces, including that an unpollable character device
such as /dev/null fails assign() on epoll and kqueue but is adopted and
works on select and io_uring; and that wait(wait_type::error) never
completes for a pipe hangup on kqueue or select, which raise no error
event for it, so wait(wait_type::read) is the uniform choice. The
reference examples for posix_descriptor::wait and the native twin are
compiled snippets under test/doc, not hand-typed code blocks, and the
file assign() return codes are scoped to the backends that produce them.
mvandeberg
force-pushed
the
pr/posix-descriptor
branch
from
September 28, 2026 19:57
6a09190 to
2c4fc5b
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
posix_descriptor adopts an already-open pollable file descriptor — a character device, inotify, eventfd, timerfd, pidfd, pipe, tty, or a socket kind corosio does not otherwise wrap — and drives it from the io_context with read_some, write_some and wait(). It never creates a descriptor; the caller supplies one. It ships on epoll, kqueue, select and io_uring, with a devirtualized native_posix_descriptor twin.
The name is platform-qualified deliberately. Portability comes from the interfaces, not the name: a posix_descriptor is an io_stream, so capy::read/write, capy::Stream-constrained algorithms and TLS layering work on it exactly as they do on a socket.
Two contracts shape the implementation:
assign() validates before it mutates. A rejected descriptor leaves the object holding whatever it held before, pending operations included, and leaves ownership of the descriptor with the caller. Regular files, block devices and directories are rejected with operation_not_supported; stream_file and random_access_file adopt those. The file-type test is a reject-list rather than an accept-list because eventfd, timerfd, inotify and pidfd are anonymous inodes whose st_mode type bits are all zero, so an accept-list would reject exactly the kinds this type exists to carry.
O_NONBLOCK is applied lazily at the first read_some/write_some, never at assign(), and never restored. wait() never touches it, so adopting STDIN_FILENO to await readiness cannot flip the parent shell's terminal to nonblocking. The flag lives on the shared open file description, which is also why restoring it would race every other holder, and why a dup() is no escape.
On io_uring the lazy flag is a cancellability requirement rather than only a contract: a blocking read punted to an io-wq worker cannot be cancelled. READV/WRITEV are submitted at offset -1 so the kernel uses the descriptor's own file position, and an EAGAIN completion arms a poll_add and resubmits. That second submission opens a window in which a cancel can land between the EAGAIN completion and its dispatch, invisible to both cancel-by-fd and a one-shot stop_callback, so the implementation carries two generation counters — one bumped by cancel(), one by any descriptor change — snapshotted at prepare time and compared before the re-arm. They are separate because the two abandonment reasons name different codes to the caller.
Three defects in shared reactor code surface through non-socket descriptors and are fixed here:
epoll delivers EPOLLHUP alone for a pipe or tty whose peer closed, and that bit mapped to no reactor event, so a parked read never ran and the edge-triggered registration never repeated: io_context::run() hung forever. EPOLLHUP now widens to read and write readiness. Sockets already carry both bits alongside EPOLLHUP, measured.
getsockopt(SO_ERROR) fails with ENOTSOCK on a non-socket, and that errno reached the caller in place of the kernel's real error. On ENOTSOCK the probe now yields, and the dispatch is forced to re-run each parked operation's own syscall so EPIPE or EIO surfaces instead.
io_uring treated POLLHUP in a poll's revents as a fault and substituted EIO, while the reactors treat it as readiness. POLLHUP is now readiness for wait(read) and wait(write); only POLLERR and POLLNVAL name an error, and wait(error) is unchanged.
stream_file::assign() and random_access_file::assign() now validate before mutating, matching the descriptor contract, and cancel in-flight work before closing the file they hold — the POSIX pool reads fd_ and offset_ at execution time, so an uncancelled operation would otherwise complete against the newly adopted file and report success.
This breaks one shipped behavior: both file types now reject a pipe, which they cannot position with preadv/pwritev, and posix_descriptor is the home for those. IOCP is unchanged; its assign() still closes the held handle before registering, and the docstrings now say so.
Adds a guide page with compiled snippets covering adoption, the dup() and descriptor-flag rules, what is rejected, and the SIGPIPE limit — a write to a descriptor whose peer has closed raises it, unlike the socket types, because MSG_NOSIGNAL is a send() flag with no writev equivalent. Two backend asymmetries are documented rather than papered over: where a kernel refusal surfaces, including that an unpollable character device such as /dev/null fails assign() on epoll and kqueue but is adopted and works on select and io_uring; and that wait(wait_type::error) never completes for a pipe hangup on kqueue or select, which raise no error event for it, so wait(wait_type::read) is the uniform choice.
Closes #338