Skip to content

feat(posix): add posix_descriptor for adopting native descriptors - #372

Open
mvandeberg wants to merge 1 commit into
cppalliance:developfrom
mvandeberg:pr/posix-descriptor
Open

mvandeberg wants to merge 1 commit into
cppalliance:developfrom
mvandeberg:pr/posix-descriptor

Conversation

@mvandeberg

@mvandeberg mvandeberg commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

posix_descriptor adopts an already-open pollable file descriptor — a character device, inotify, eventfd, timerfd, pidfd, pipe, tty, or a socket kind corosio does not otherwise wrap — and drives it from the io_context with read_some, write_some and wait(). It never creates a descriptor; the caller supplies one. It ships on epoll, kqueue, select and io_uring, with a devirtualized native_posix_descriptor twin.

The name is platform-qualified deliberately. Portability comes from the interfaces, not the name: a posix_descriptor is an io_stream, so capy::read/write, capy::Stream-constrained algorithms and TLS layering work on it exactly as they do on a socket.

Two contracts shape the implementation:

  • assign() validates before it mutates. A rejected descriptor leaves the object holding whatever it held before, pending operations included, and leaves ownership of the descriptor with the caller. Regular files, block devices and directories are rejected with operation_not_supported; stream_file and random_access_file adopt those. The file-type test is a reject-list rather than an accept-list because eventfd, timerfd, inotify and pidfd are anonymous inodes whose st_mode type bits are all zero, so an accept-list would reject exactly the kinds this type exists to carry.

  • O_NONBLOCK is applied lazily at the first read_some/write_some, never at assign(), and never restored. wait() never touches it, so adopting STDIN_FILENO to await readiness cannot flip the parent shell's terminal to nonblocking. The flag lives on the shared open file description, which is also why restoring it would race every other holder, and why a dup() is no escape.

On io_uring the lazy flag is a cancellability requirement rather than only a contract: a blocking read punted to an io-wq worker cannot be cancelled. READV/WRITEV are submitted at offset -1 so the kernel uses the descriptor's own file position, and an EAGAIN completion arms a poll_add and resubmits. That second submission opens a window in which a cancel can land between the EAGAIN completion and its dispatch, invisible to both cancel-by-fd and a one-shot stop_callback, so the implementation carries two generation counters — one bumped by cancel(), one by any descriptor change — snapshotted at prepare time and compared before the re-arm. They are separate because the two abandonment reasons name different codes to the caller.

Three defects in shared reactor code surface through non-socket descriptors and are fixed here:

  • epoll delivers EPOLLHUP alone for a pipe or tty whose peer closed, and that bit mapped to no reactor event, so a parked read never ran and the edge-triggered registration never repeated: io_context::run() hung forever. EPOLLHUP now widens to read and write readiness. Sockets already carry both bits alongside EPOLLHUP, measured.

  • getsockopt(SO_ERROR) fails with ENOTSOCK on a non-socket, and that errno reached the caller in place of the kernel's real error. On ENOTSOCK the probe now yields, and the dispatch is forced to re-run each parked operation's own syscall so EPIPE or EIO surfaces instead.

  • io_uring treated POLLHUP in a poll's revents as a fault and substituted EIO, while the reactors treat it as readiness. POLLHUP is now readiness for wait(read) and wait(write); only POLLERR and POLLNVAL name an error, and wait(error) is unchanged.

stream_file::assign() and random_access_file::assign() now validate before mutating, matching the descriptor contract, and cancel in-flight work before closing the file they hold — the POSIX pool reads fd_ and offset_ at execution time, so an uncancelled operation would otherwise complete against the newly adopted file and report success.

This breaks one shipped behavior: both file types now reject a pipe, which they cannot position with preadv/pwritev, and posix_descriptor is the home for those. IOCP is unchanged; its assign() still closes the held handle before registering, and the docstrings now say so.

Adds a guide page with compiled snippets covering adoption, the dup() and descriptor-flag rules, what is rejected, and the SIGPIPE limit — a write to a descriptor whose peer has closed raises it, unlike the socket types, because MSG_NOSIGNAL is a send() flag with no writev equivalent. Two backend asymmetries are documented rather than papered over: where a kernel refusal surfaces, including that an unpollable character device such as /dev/null fails assign() on epoll and kqueue but is adopted and works on select and io_uring; and that wait(wait_type::error) never completes for a pipe hangup on kqueue or select, which raise no error event for it, so wait(wait_type::read) is the uniform choice.

Closes #338

@cppalliance-bot

cppalliance-bot commented Sep 28, 2026 •

Copy link
Copy Markdown

An automated preview of the documentation is available at https://372.corosio.prtest3.cppalliance.org/index.html

If more commits are pushed to the pull request, the docs will rebuild at the same URL.

2026-09-28 19:59:33 UTC

@cppalliance-bot

cppalliance-bot commented Sep 28, 2026 •

Copy link
Copy Markdown

GCOVR code coverage report https://372.corosio.prtest3.cppalliance.org/gcovr/index.html
LCOV code coverage report https://372.corosio.prtest3.cppalliance.org/genhtml/index.html
Coverage Diff Report https://372.corosio.prtest3.cppalliance.org/diff-report/index.html

Build time: 2026-09-28 20:15:35 UTC

@mvandeberg
mvandeberg force-pushed the pr/posix-descriptor branch 2 times, most recently from 92bcedf to 6a09190 Compare September 28, 2026 19:08
posix_descriptor adopts an already-open pollable file descriptor — a
character device, inotify, eventfd, timerfd, pidfd, pipe, tty, or a socket
kind corosio does not otherwise wrap — and drives it from the io_context
with read_some, write_some and wait(). It never creates a descriptor; the
caller supplies one. It ships on epoll, kqueue, select and io_uring, with
a devirtualized native_posix_descriptor<Backend> twin.

The name is platform-qualified deliberately. Portability comes from the
interfaces, not the name: a posix_descriptor is an io_stream, so
capy::read/write, capy::Stream-constrained algorithms and TLS layering
work on it exactly as they do on a socket.

Two contracts shape the implementation:

  - assign() validates before it mutates. A rejected descriptor leaves the
    object holding whatever it held before, pending operations included,
    and leaves ownership of the descriptor with the caller. Regular files,
    block devices and directories are rejected with
    operation_not_supported; stream_file and random_access_file adopt
    those. The file-type test is a reject-list rather than an accept-list
    because eventfd, timerfd, inotify and pidfd are anonymous inodes whose
    st_mode type bits are all zero, so an accept-list would reject exactly
    the kinds this type exists to carry.

  - O_NONBLOCK is applied lazily at the first read_some/write_some, never
    at assign(), and never restored. wait() never touches it, so adopting
    STDIN_FILENO to await readiness cannot flip the parent shell's
    terminal to nonblocking. The flag lives on the shared open file
    description, which is also why restoring it would race every other
    holder, and why a dup() is no escape.

On io_uring the lazy flag is a cancellability requirement rather than only
a contract: a blocking read punted to an io-wq worker cannot be cancelled.
READV/WRITEV are submitted at offset -1 so the kernel uses the
descriptor's own file position, and an EAGAIN completion arms a poll_add
and resubmits. That second submission opens a window in which a cancel can
land between the EAGAIN completion and its dispatch, invisible to both
cancel-by-fd and a one-shot stop_callback, so the implementation carries
two generation counters — one bumped by cancel(), one by any descriptor
change — snapshotted at prepare time and compared before the re-arm. They
are separate because the two abandonment reasons name different codes to
the caller.

Three defects in shared reactor code surface through non-socket
descriptors and are fixed here:

  - epoll delivers EPOLLHUP alone for a pipe or tty whose peer closed, and
    that bit mapped to no reactor event, so a parked read never ran and
    the edge-triggered registration never repeated: io_context::run() hung
    forever. EPOLLHUP now widens to read and write readiness. Sockets
    already carry both bits alongside EPOLLHUP, measured.

  - getsockopt(SO_ERROR) fails with ENOTSOCK on a non-socket, and that
    errno reached the caller in place of the kernel's real error. On
    ENOTSOCK the probe now yields, and the dispatch is forced to re-run
    each parked operation's own syscall so EPIPE or EIO surfaces instead.

  - io_uring treated POLLHUP in a poll's revents as a fault and
    substituted EIO, while the reactors treat it as readiness. POLLHUP is
    now readiness for wait(read) and wait(write); only POLLERR and
    POLLNVAL name an error, and wait(error) is unchanged.

The devirtualized twin's read_some, write_some and wait awaitables
hand-rolled an await_ready/await_resume pair that predates op_base, testing
the stop token directly. That reported a cancellation on an operation the
backend had already completed successfully. All three now derive from
bytes_op_base and void_op_base like every other awaitable, so a stop request
arriving after completion no longer displaces the result.

stream_file::assign() and random_access_file::assign() now validate before
mutating, matching the descriptor contract, and cancel in-flight work
before closing the file they hold — the POSIX pool reads fd_ and offset_
at execution time, so an uncancelled operation would otherwise complete
against the newly adopted file and report success.

This breaks one shipped behavior: both file types now reject a pipe, which
they cannot position with preadv/pwritev, and posix_descriptor is the home
for those. IOCP is unchanged; its assign() still closes the held handle
before registering, and the docstrings now say so.

Adds a guide page with compiled snippets covering adoption, the dup() and
descriptor-flag rules, what is rejected, and the SIGPIPE limit — a write
to a descriptor whose peer has closed raises it, unlike the socket types,
because MSG_NOSIGNAL is a send() flag with no writev equivalent. Two
backend asymmetries are documented rather than papered over: where a
kernel refusal surfaces, including that an unpollable character device
such as /dev/null fails assign() on epoll and kqueue but is adopted and
works on select and io_uring; and that wait(wait_type::error) never
completes for a pipe hangup on kqueue or select, which raise no error
event for it, so wait(wait_type::read) is the uniform choice. The
reference examples for posix_descriptor::wait and the native twin are
compiled snippets under test/doc, not hand-typed code blocks, and the
file assign() return codes are scoped to the backends that produce them.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

Native descriptor and handle support: posix_descriptor, win_stream_handle, win_random_access_handle, win_object_handle

2 participants