From 2c4fc5b802563a74da24af20cf21d514affb8ea1 Mon Sep 17 00:00:00 2001 From: Michael Vandeberg Date: Mon, 28 Sep 2026 12:00:51 -0600 Subject: [PATCH] feat(posix): add posix_descriptor for adopting native descriptors MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit posix_descriptor adopts an already-open pollable file descriptor — a character device, inotify, eventfd, timerfd, pidfd, pipe, tty, or a socket kind corosio does not otherwise wrap — and drives it from the io_context with read_some, write_some and wait(). It never creates a descriptor; the caller supplies one. It ships on epoll, kqueue, select and io_uring, with a devirtualized native_posix_descriptor twin. The name is platform-qualified deliberately. Portability comes from the interfaces, not the name: a posix_descriptor is an io_stream, so capy::read/write, capy::Stream-constrained algorithms and TLS layering work on it exactly as they do on a socket. Two contracts shape the implementation: - assign() validates before it mutates. A rejected descriptor leaves the object holding whatever it held before, pending operations included, and leaves ownership of the descriptor with the caller. Regular files, block devices and directories are rejected with operation_not_supported; stream_file and random_access_file adopt those. The file-type test is a reject-list rather than an accept-list because eventfd, timerfd, inotify and pidfd are anonymous inodes whose st_mode type bits are all zero, so an accept-list would reject exactly the kinds this type exists to carry. - O_NONBLOCK is applied lazily at the first read_some/write_some, never at assign(), and never restored. wait() never touches it, so adopting STDIN_FILENO to await readiness cannot flip the parent shell's terminal to nonblocking. The flag lives on the shared open file description, which is also why restoring it would race every other holder, and why a dup() is no escape. On io_uring the lazy flag is a cancellability requirement rather than only a contract: a blocking read punted to an io-wq worker cannot be cancelled. READV/WRITEV are submitted at offset -1 so the kernel uses the descriptor's own file position, and an EAGAIN completion arms a poll_add and resubmits. That second submission opens a window in which a cancel can land between the EAGAIN completion and its dispatch, invisible to both cancel-by-fd and a one-shot stop_callback, so the implementation carries two generation counters — one bumped by cancel(), one by any descriptor change — snapshotted at prepare time and compared before the re-arm. They are separate because the two abandonment reasons name different codes to the caller. Three defects in shared reactor code surface through non-socket descriptors and are fixed here: - epoll delivers EPOLLHUP alone for a pipe or tty whose peer closed, and that bit mapped to no reactor event, so a parked read never ran and the edge-triggered registration never repeated: io_context::run() hung forever. EPOLLHUP now widens to read and write readiness. Sockets already carry both bits alongside EPOLLHUP, measured. - getsockopt(SO_ERROR) fails with ENOTSOCK on a non-socket, and that errno reached the caller in place of the kernel's real error. On ENOTSOCK the probe now yields, and the dispatch is forced to re-run each parked operation's own syscall so EPIPE or EIO surfaces instead. - io_uring treated POLLHUP in a poll's revents as a fault and substituted EIO, while the reactors treat it as readiness. POLLHUP is now readiness for wait(read) and wait(write); only POLLERR and POLLNVAL name an error, and wait(error) is unchanged. The devirtualized twin's read_some, write_some and wait awaitables hand-rolled an await_ready/await_resume pair that predates op_base, testing the stop token directly. That reported a cancellation on an operation the backend had already completed successfully. All three now derive from bytes_op_base and void_op_base like every other awaitable, so a stop request arriving after completion no longer displaces the result. stream_file::assign() and random_access_file::assign() now validate before mutating, matching the descriptor contract, and cancel in-flight work before closing the file they hold — the POSIX pool reads fd_ and offset_ at execution time, so an uncancelled operation would otherwise complete against the newly adopted file and report success. This breaks one shipped behavior: both file types now reject a pipe, which they cannot position with preadv/pwritev, and posix_descriptor is the home for those. IOCP is unchanged; its assign() still closes the held handle before registering, and the docstrings now say so. Adds a guide page with compiled snippets covering adoption, the dup() and descriptor-flag rules, what is rejected, and the SIGPIPE limit — a write to a descriptor whose peer has closed raises it, unlike the socket types, because MSG_NOSIGNAL is a send() flag with no writev equivalent. Two backend asymmetries are documented rather than papered over: where a kernel refusal surfaces, including that an unpollable character device such as /dev/null fails assign() on epoll and kqueue but is adopted and works on select and io_uring; and that wait(wait_type::error) never completes for a pipe hangup on kqueue or select, which raise no error event for it, so wait(wait_type::read) is the uniform choice. The reference examples for posix_descriptor::wait and the native twin are compiled snippets under test/doc, not hand-typed code blocks, and the file assign() return codes are scoped to the backends that produce them. --- .../config/vocabularies/Corosio/accept.txt | 4 + doc/modules/ROOT/nav.adoc | 1 + .../ROOT/pages/4.guide/4o.file-io.adoc | 16 +- doc/modules/ROOT/pages/4.guide/4r.wait.adoc | 7 + .../pages/4.guide/4s.native-descriptors.adoc | 224 +++++ include/boost/corosio.hpp | 2 + include/boost/corosio/backend.hpp | 29 + .../corosio/detail/descriptor_service.hpp | 66 ++ .../native/detail/epoll/epoll_scheduler.hpp | 24 +- .../native/detail/epoll/epoll_traits.hpp | 35 + .../native/detail/epoll/epoll_types.hpp | 44 + .../native/detail/kqueue/kqueue_traits.hpp | 33 + .../native/detail/kqueue/kqueue_types.hpp | 44 + .../detail/posix/posix_random_access_file.hpp | 18 + .../native/detail/posix/posix_stream_file.hpp | 18 + .../detail/reactor/reactor_descriptor.hpp | 872 ++++++++++++++++++ .../reactor/reactor_descriptor_service.hpp | 171 ++++ .../reactor/reactor_descriptor_state.hpp | 27 +- .../native/detail/select/select_traits.hpp | 34 + .../native/detail/select/select_types.hpp | 44 + .../native/detail/uring/uring_descriptor.hpp | 569 ++++++++++++ .../detail/uring/uring_descriptor_service.hpp | 143 +++ .../detail/uring/uring_random_access_file.hpp | 16 + .../native/detail/uring/uring_socket_ops.hpp | 26 +- .../native/detail/uring/uring_stream_file.hpp | 16 + .../corosio/native/detail/validate_fd.hpp | 111 +++ include/boost/corosio/native/native.hpp | 2 + .../native/native_posix_descriptor.hpp | 229 +++++ include/boost/corosio/posix_descriptor.hpp | 362 ++++++++ include/boost/corosio/random_access_file.hpp | 40 +- include/boost/corosio/stream_file.hpp | 42 +- .../src/detail/use_backend_service.hpp | 7 + src/corosio/src/posix_descriptor.cpp | 80 ++ src/corosio/src/random_access_file.cpp | 5 +- src/corosio/src/stream_file.cpp | 5 +- .../native_posix_descriptor.record.cpp | 69 ++ .../posix_descriptor__wait.function.cpp | 62 ++ test/doc/snippets/4s_native_descriptors.cpp | 273 ++++++ test/unit/lazy_services.cpp | 9 + test/unit/native/native_posix_descriptor.cpp | 156 ++++ test/unit/native/native_resume_cancel.cpp | 95 ++ .../unit/native/uring/descriptor_continue.cpp | 302 ++++++ test/unit/native/validate_fd.cpp | 169 ++++ test/unit/posix_descriptor.cpp | 796 ++++++++++++++++ test/unit/random_access_file.cpp | 134 ++- test/unit/stream_file.cpp | 138 ++- test/unit/teardown_inflight.cpp | 57 +- 47 files changed, 5567 insertions(+), 59 deletions(-) create mode 100644 doc/modules/ROOT/pages/4.guide/4s.native-descriptors.adoc create mode 100644 include/boost/corosio/detail/descriptor_service.hpp create mode 100644 include/boost/corosio/native/detail/reactor/reactor_descriptor.hpp create mode 100644 include/boost/corosio/native/detail/reactor/reactor_descriptor_service.hpp create mode 100644 include/boost/corosio/native/detail/uring/uring_descriptor.hpp create mode 100644 include/boost/corosio/native/detail/uring/uring_descriptor_service.hpp create mode 100644 include/boost/corosio/native/native_posix_descriptor.hpp create mode 100644 include/boost/corosio/posix_descriptor.hpp create mode 100644 src/corosio/src/posix_descriptor.cpp create mode 100644 test/doc/reference/native_posix_descriptor.record.cpp create mode 100644 test/doc/reference/posix_descriptor__wait.function.cpp create mode 100644 test/doc/snippets/4s_native_descriptors.cpp create mode 100644 test/unit/native/native_posix_descriptor.cpp create mode 100644 test/unit/native/uring/descriptor_continue.cpp create mode 100644 test/unit/native/validate_fd.cpp create mode 100644 test/unit/posix_descriptor.cpp diff --git a/doc/.vale/styles/config/vocabularies/Corosio/accept.txt b/doc/.vale/styles/config/vocabularies/Corosio/accept.txt index 3275f5036..631a2772a 100644 --- a/doc/.vale/styles/config/vocabularies/Corosio/accept.txt +++ b/doc/.vale/styles/config/vocabularies/Corosio/accept.txt @@ -89,6 +89,8 @@ (?i)kqueue (?i)io_uring (?i)iovec +(?i)ttys? +(?i)inodes? (?i)syscalls? (?i)datagrams? (?i)wakeups? @@ -133,6 +135,7 @@ # --- Coined adjectives and nouns the guide uses --------------------------------- (?i)joinable +(?i)pollable (?i)schedulable (?i)launchable (?i)buildable @@ -278,6 +281,7 @@ (?i)wolfssl (?i)backend[’']s (?i)scheduler[’']s +(?i)select[’']s (?i)IP[’']s (?i)UDP[’']s (?i)URL[’']s diff --git a/doc/modules/ROOT/nav.adoc b/doc/modules/ROOT/nav.adoc index 2005b0d12..b52330919 100644 --- a/doc/modules/ROOT/nav.adoc +++ b/doc/modules/ROOT/nav.adoc @@ -40,6 +40,7 @@ ** xref:4.guide/4p.unix-sockets.adoc[Unix Domain Sockets] ** xref:4.guide/4q.udp.adoc[UDP Sockets] ** xref:4.guide/4r.wait.adoc[Readiness Wait] +** xref:4.guide/4s.native-descriptors.adoc[Native Descriptors] * xref:5.testing/5.intro.adoc[Testing] ** xref:5.testing/5a.mocket.adoc[Mock Sockets] ** xref:5.testing/5b.socket-pair.adoc[Socket Pairs] diff --git a/doc/modules/ROOT/pages/4.guide/4o.file-io.adoc b/doc/modules/ROOT/pages/4.guide/4o.file-io.adoc index dbbd52107..93b8d0247 100644 --- a/doc/modules/ROOT/pages/4.guide/4o.file-io.adoc +++ b/doc/modules/ROOT/pages/4.guide/4o.file-io.adoc @@ -88,7 +88,7 @@ Both file types accept a bitmask of `file_base::flags` when opening: | `create` | Create the file if it does not exist | `exclusive` | Fail if the file already exists (requires `create`) | `truncate` | Truncate the file to zero length on open -| `append` | Seek to end on open (stream_file only) +| `append` | Seek to end on open (cpp:stream_file[] only) | `sync_all_on_write` | Synchronize data to disk on each write |=== @@ -137,6 +137,20 @@ file object is already associated and cannot be re-adopted there. On POSIX platforms no such restriction exists. ==== +On POSIX, `assign()` accepts only what a file object can position: +regular files, block devices, and character devices. A pipe or socket is +rejected with `errc::operation_not_supported` — adopt those into a +xref:4.guide/4s.native-descriptors.adoc[`posix_descriptor`] instead; a +directory is not adoptable by either type. + +A character device that cannot seek, such as a tty, passes adoption on +every backend. What happens next is backend-specific: the POSIX +backends issue `preadv`/`pwritev` and fail at the first read or write +with `ESPIPE`. io_uring submits `READV`/`WRITEV` at offset `-1` for a +`stream_file` and reads the tty successfully. Adopt one into a +xref:4.guide/4s.native-descriptors.adoc[`posix_descriptor`] if you want +the same behavior everywhere. + == Error Handling File operations follow the same error model as sockets. Reads past diff --git a/doc/modules/ROOT/pages/4.guide/4r.wait.adoc b/doc/modules/ROOT/pages/4.guide/4r.wait.adoc index f1bb0c514..8b240a82b 100644 --- a/doc/modules/ROOT/pages/4.guide/4r.wait.adoc +++ b/doc/modules/ROOT/pages/4.guide/4r.wait.adoc @@ -70,6 +70,12 @@ library's socket directly and `release()` it before the library needs exclusive ownership again. Or create a true duplicate with `WSADuplicateSocketW` and adopt that. +This applies to sockets. A library that owns a non-socket descriptor — +a pipe, a tty, `inotify`, `eventfd` — needs the same readiness +notification. +xref:4.guide/4s.native-descriptors.adoc[`posix_descriptor`] provides it +with the same `wait()` on the same terms. + == Acceptors cpp:tcp_acceptor[] and cpp:local_stream_acceptor[] expose the same `wait()`. @@ -150,3 +156,4 @@ uniform across platforms. * xref:4.guide/4d.sockets.adoc[Sockets] * xref:4.guide/4e.tcp-acceptor.adoc[Acceptors] * xref:4.guide/4q.udp.adoc[UDP Sockets] +* xref:4.guide/4s.native-descriptors.adoc[Native Descriptors] diff --git a/doc/modules/ROOT/pages/4.guide/4s.native-descriptors.adoc b/doc/modules/ROOT/pages/4.guide/4s.native-descriptors.adoc new file mode 100644 index 000000000..6c30e41da --- /dev/null +++ b/doc/modules/ROOT/pages/4.guide/4s.native-descriptors.adoc @@ -0,0 +1,224 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + += Native Descriptors +:page-mode: how-to + +cpp:posix_descriptor[] adopts a file descriptor you already have and +drives it from an `io_context`, giving it `read_some()`, `write_some()` +and `wait()`. It exists on POSIX platforms only. + +[NOTE] +==== +Code snippets assume: +[source,cpp] +---- +include::example$snippets/4s_native_descriptors.cpp[tag=assume] +---- +==== + +== Overview + +Corosio calls a descriptor pollable when a reactor can wait on it for +readiness. Anything pollable that corosio does not already wrap is in +scope. That includes character devices, `inotify`, `eventfd`, +`timerfd`, `pidfd`, pipes, ttys, and socket kinds that have no +dedicated corosio type. + +The type never creates a descriptor. Open it with whichever platform +call suits it — `eventfd()`, `inotify_init1()`, `open()` on a device +node — and hand the result to `assign()`. corosio supplies the event +loop, not the constructor. + +== Why the Name Says POSIX + +A single portable `native_descriptor` spanning POSIX and Windows was +considered and rejected. The platform gaps here are not edge cases +around a shared core; they are the type's semantics. `O_NONBLOCK` on a +shared open file description, `dup()`, and a file-type reject list +expressed in `st_mode` bits are the entire contract below. An open +file description is the kernel-side object a descriptor refers to. +Every `dup()` of a descriptor shares the same one. None of them has a +Windows counterpart. A type that named both would have to either +document each rule twice or say nothing precise about either. + +Portability lives one layer up instead. A cpp:posix_descriptor[] is an +cpp:io_object[], cpp:io_read_stream[], cpp:io_write_stream[] and +cpp:io_stream[] — the same bases `tcp_socket` has — and, like every +corosio stream, it satisfies `capy::Stream`. That concept is what the +generic algorithms are written against, so they run on a descriptor, a +socket and a `tls_stream` alike: + +[source,cpp] +---- +include::example$snippets/4s_native_descriptors.cpp[tag=layering,indent=0] +---- + +Only the handful of lines that produce the descriptor are +platform-specific. xref:4.guide/4l.tls.adoc[TLS] layers over it on the +same terms. + +== Adopting a Descriptor + +`assign()` takes ownership: `close()` and the destructor close the +descriptor. It returns a `std::error_code` and is pass:[[[nodiscard]]]; a +failed `assign()` leaves the descriptor with you, so close it yourself. + +An `eventfd` as a cross-thread wakeup: + +[source,cpp] +---- +include::example$snippets/4s_native_descriptors.cpp[tag=adopt_eventfd,indent=0] +---- + +[NOTE] +==== +Distinct cpp:posix_descriptor[] objects are safe to use from different +threads. A shared object must not run two operations of the same kind +at once. One read and one write may overlap. +==== + +An `inotify` watch. The descriptor is a stream of variable-length +records, so `read_some()` is the whole interface you need: + +[source,cpp] +---- +include::example$snippets/4s_native_descriptors.cpp[tag=adopt_inotify,indent=0] +---- + +`release()` hands the descriptor back, cancelling pending operations +and leaving the object not-open. + +== Ownership and the `dup()` Rule + +Another party sometimes owns the descriptor — a C library that does +its own I/O on it, or a process-wide descriptor such as +`STDIN_FILENO`. When that happens, adopt a `dup()` of it rather than +the descriptor itself: + +[source,cpp] +---- +include::example$snippets/4s_native_descriptors.cpp[tag=dup_for_foreign_fd,indent=0] +---- + +Both descriptors refer to one open file description, so the duplicate +reports exactly the original's readiness. Corosio closing the +duplicate can never close the original. This is the same rule +xref:4.guide/4r.wait.adoc[Readiness Wait] states for adopted sockets. + +== Descriptor Flags + +[WARNING] +==== +`O_NONBLOCK` is set on the first `read_some()` or `write_some()`, never +by `assign()`, and it is never restored. + +The flag lives on the shared open file description, not on the +descriptor, so every other holder of that description sees it. +Restoring it on close would race whoever else is holding it. Permanent +is the only safe choice. + +A `dup()` does not shield the other holder from this: the duplicate +shares the same description, so the flag change reaches them anyway. It +separates the lifetimes, nothing more. When another party owns the +descriptor and cannot tolerate `O_NONBLOCK`, the way out is `wait()` +and doing the I/O yourself. `wait()` never modifies the descriptor at +all, flags included. +==== + +That is what makes standard input safe to adopt for readiness alone. +Flipping `O_NONBLOCK` on it would change the terminal the parent shell +is still using. + +[source,cpp] +---- +include::example$snippets/4s_native_descriptors.cpp[tag=wait_only,indent=0] +---- + +== What Is Rejected + +`assign()` rejects regular files, block devices, and directories with a +code comparing equal to `errc::operation_not_supported`. A reactor +cannot report readiness for them, and they already have a home: +xref:4.guide/4o.file-io.adoc[`stream_file` and `random_access_file`] +adopt exactly those kinds. + +The test is a reject list, not an accept list. The reason: the +flagship descriptor kinds — `eventfd`, `timerfd`, `inotify`, `pidfd` — +are anonymous inodes whose `st_mode` type bits are all zero. An accept list +would reject the descriptors this type exists to carry. + +A negative or closed descriptor fails with `errc::bad_file_descriptor`, +and re-assigning the descriptor the object already holds fails with +`errc::invalid_argument`. + +== Where Errors Surface + +Validation runs before anything is mutated. A rejected descriptor +leaves the object holding whatever it held before — pending operations +included — and leaves you owning the descriptor. + +A refusal from the *kernel* is the one exception, and where it appears +depends on the backend: + +epoll, kqueue, select:: These register the descriptor with the reactor +during `assign()`, so a refusal fails `assign()`. The old descriptor +has already been closed by then, so the object is left closed. On +select this is reachable in normal use: `select()` cannot monitor a +descriptor at or above `FD_SETSIZE`, and such a descriptor is rejected +with `EMFILE`. + +io_uring:: There is no adopt-time registration syscall, so `assign()` +succeeds and a refusal appears at the first operation instead. + +Whichever way `assign()` fails, the descriptor you passed is still +yours to close. + +Not every character device can be adopted. `/dev/null`, `/dev/zero` and +`/dev/urandom` are not pollable: epoll refuses them with `EPERM` and +kqueue with `EINVAL`, so `assign()` fails on those backends. select and +io_uring have nothing to refuse them with, so `assign()` succeeds and the +descriptor works — a `read_some()` on `/dev/zero` returns zeros. Character +devices backed by a real driver, a tty among them, are pollable and +adopt everywhere. + +`wait(wait_type::error)` is the one verb that is not uniform. epoll and +io_uring report a pipe or FIFO hangup as an error condition and name a +code. kqueue and select do not. kqueue raises an error event only for +`EV_ERROR` or for `EV_EOF` with `fflags != 0`, and a hangup sets +neither. select's exceptional set does not cover it. On those two backends +the wait never completes; end it with `cancel()` or a stop token. Prefer +`wait(wait_type::read)`, which is uniform — the hangup surfaces there as +readiness, and the read that follows names the real failure. + +== `SIGPIPE` + +[WARNING] +==== +Writing to a descriptor whose peer has closed raises `SIGPIPE` in the +default disposition, which terminates the process. The socket types +suppress this; cpp:posix_descriptor[] cannot. + +The suppression sockets get has no general form. `MSG_NOSIGNAL` is a +`send()` flag and there is no `writev()` equivalent. `SO_NOSIGPIPE` +is a socket option. Neither applies to an arbitrary descriptor. + +Install `SIG_IGN` for `SIGPIPE` — or handle it through a +xref:4.guide/4i.signals.adoc[`signal_set`] — before writing to an +adopted descriptor. The write then fails with `EPIPE` instead. +==== + +Asio's `posix::stream_descriptor` behaves the same way, for the same +reason. Code ported from it needs no change here. + +== See Also + +* xref:4.guide/4r.wait.adoc[Readiness Wait] +* xref:4.guide/4o.file-io.adoc[File I/O] +* xref:reference:boost/corosio/posix_descriptor.adoc[`posix_descriptor` reference] diff --git a/include/boost/corosio.hpp b/include/boost/corosio.hpp index a6bcb41aa..4337cc84e 100644 --- a/include/boost/corosio.hpp +++ b/include/boost/corosio.hpp @@ -1,5 +1,6 @@ // // Copyright (c) 2025 Vinnie Falco (vinnie.falco@gmail.com) +// Copyright (c) 2026 Michael Vandeberg // // Distributed under the Boost Software License, Version 1.0. (See accompanying // file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) @@ -22,6 +23,7 @@ #include #include #include +#include // POSIX-only; self-guarded #include #include #include diff --git a/include/boost/corosio/backend.hpp b/include/boost/corosio/backend.hpp index 1bfdbac39..1cb1caf07 100644 --- a/include/boost/corosio/backend.hpp +++ b/include/boost/corosio/backend.hpp @@ -1,5 +1,6 @@ // // Copyright (c) 2026 Steve Gerbino +// Copyright (c) 2026 Michael Vandeberg // // Distributed under the Boost Software License, Version 1.0. (See accompanying // file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) @@ -39,6 +40,8 @@ class epoll_local_stream_acceptor; class epoll_local_stream_acceptor_service; class epoll_local_datagram_socket; class epoll_local_datagram_service; +class epoll_descriptor; +class epoll_descriptor_service; class epoll_scheduler; class posix_signal; @@ -84,6 +87,11 @@ struct epoll_t /// The service that owns the Unix domain datagram implementations. using local_datagram_service_type = detail::epoll_local_datagram_service; + /// The concrete adopted-descriptor type. + using descriptor_type = detail::epoll_descriptor; + /// The service that owns the descriptor implementations. + using descriptor_service_type = detail::epoll_descriptor_service; + /// The concrete signal set type. using signal_type = detail::posix_signal; /// The service that owns the signal set implementations. @@ -141,6 +149,8 @@ class select_local_stream_acceptor; class select_local_stream_acceptor_service; class select_local_datagram_socket; class select_local_datagram_service; +class select_descriptor; +class select_descriptor_service; class select_scheduler; class posix_signal; @@ -186,6 +196,11 @@ struct select_t /// The service that owns the Unix domain datagram implementations. using local_datagram_service_type = detail::select_local_datagram_service; + /// The concrete adopted-descriptor type. + using descriptor_type = detail::select_descriptor; + /// The service that owns the descriptor implementations. + using descriptor_service_type = detail::select_descriptor_service; + /// The concrete signal set type. using signal_type = detail::posix_signal; /// The service that owns the signal set implementations. @@ -243,6 +258,8 @@ class kqueue_local_stream_acceptor; class kqueue_local_stream_acceptor_service; class kqueue_local_datagram_socket; class kqueue_local_datagram_service; +class kqueue_descriptor; +class kqueue_descriptor_service; class kqueue_scheduler; class posix_signal; @@ -288,6 +305,11 @@ struct kqueue_t /// The service that owns the Unix domain datagram implementations. using local_datagram_service_type = detail::kqueue_local_datagram_service; + /// The concrete adopted-descriptor type. + using descriptor_type = detail::kqueue_descriptor; + /// The service that owns the descriptor implementations. + using descriptor_service_type = detail::kqueue_descriptor_service; + /// The concrete signal set type. using signal_type = detail::posix_signal; /// The service that owns the signal set implementations. @@ -345,6 +367,8 @@ class uring_local_stream_acceptor; class uring_local_stream_acceptor_service; class uring_local_datagram_socket; class uring_local_datagram_service; +class uring_descriptor; +class uring_descriptor_service; class uring_stream_file; class uring_stream_file_service; class uring_random_access_file; @@ -390,6 +414,11 @@ struct uring_t /// The service that owns the Unix domain datagram implementations. using local_datagram_service_type = detail::uring_local_datagram_service; + /// The concrete adopted-descriptor type. + using descriptor_type = detail::uring_descriptor; + /// The service that owns the descriptor implementations. + using descriptor_service_type = detail::uring_descriptor_service; + /// The concrete signal set type. using signal_type = detail::posix_signal; /// The service that owns the signal set implementations. diff --git a/include/boost/corosio/detail/descriptor_service.hpp b/include/boost/corosio/detail/descriptor_service.hpp new file mode 100644 index 000000000..b3aa4da28 --- /dev/null +++ b/include/boost/corosio/detail/descriptor_service.hpp @@ -0,0 +1,66 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +#ifndef BOOST_COROSIO_DETAIL_DESCRIPTOR_SERVICE_HPP +#define BOOST_COROSIO_DETAIL_DESCRIPTOR_SERVICE_HPP + +#include +#include + +#if BOOST_COROSIO_POSIX + +#include +#include + +#include + +namespace boost::corosio::detail { + +/* Abstract descriptor service base class. + + Concrete implementations (epoll, select, kqueue, uring) inherit + from this class. The three reactor backends register the adopted + fd with their reactor; uring has no adopt-time registration and + only takes ownership of it. + The context constructor installs whichever backend via + make_service, and posix_descriptor.cpp retrieves it via + create_handle(). +*/ +class BOOST_COROSIO_DECL descriptor_service + : public capy::execution_context::service + , public io_object::io_service +{ +public: + /// Identifies this service for execution_context lookup. + using key_type = descriptor_service; + + /** Adopt an existing native descriptor. + + Validates before mutating: on failure the implementation + keeps its previous descriptor and pending operations, and + the caller retains ownership of @a fd. On success the + implementation takes ownership and will close it. + + @param impl The descriptor implementation to assign to. + @param fd The native descriptor to adopt. + @return Error code on failure, empty on success. + */ + virtual std::error_code assign_descriptor( + posix_descriptor::implementation& impl, native_handle_type fd) = 0; + +protected: + descriptor_service() = default; + ~descriptor_service() override = default; +}; + +} // namespace boost::corosio::detail + +#endif // BOOST_COROSIO_POSIX + +#endif diff --git a/include/boost/corosio/native/detail/epoll/epoll_scheduler.hpp b/include/boost/corosio/native/detail/epoll/epoll_scheduler.hpp index 574c25665..b8d05a934 100644 --- a/include/boost/corosio/native/detail/epoll/epoll_scheduler.hpp +++ b/include/boost/corosio/native/detail/epoll/epoll_scheduler.hpp @@ -387,7 +387,29 @@ epoll_scheduler::run_task(lock_type& lock, context_type& ctx, long timeout_us) auto* desc = static_cast(event_buffer_[i].data.ptr); - desc->add_ready_events(event_buffer_[i].events); + + // A pipe or tty whose peer closed reports EPOLLHUP on its own -- + // no EPOLLIN, no EPOLLERR -- and EPOLLHUP maps to no + // reactor_event_* bit, so invoke_deferred_io() would take no + // branch and, the registration being edge-triggered, never get + // another chance. + // + // Sockets are unaffected because they never report EPOLLHUP + // alone. tcp_poll() and unix_poll() raise it only once + // sk_shutdown is SHUTDOWN_MASK (or the state is TCP_CLOSE), and + // both also report the socket readable and writable there -- + // tcp_poll() takes an explicit `else mask |= EPOLLOUT` branch + // once SEND_SHUTDOWN is set, because a send on a shut-down + // socket fails fast rather than blocking. Measured on every + // state that produces EPOLLHUP -- peer close plus local + // SHUT_WR/SHUT_RDWR, RST, RST with the send buffer full, and + // the AF_UNIX equivalents -- the mask is always IN|OUT|HUP + // (0x15), or IN|OUT|ERR|HUP (0x1d) for a reset. The forced bits + // are therefore already set on every socket path. + std::uint32_t ev = event_buffer_[i].events; + if (ev & EPOLLHUP) + ev |= EPOLLIN | EPOLLOUT; + desc->add_ready_events(ev); bool expected = false; if (desc->is_enqueued_.compare_exchange_strong( diff --git a/include/boost/corosio/native/detail/epoll/epoll_traits.hpp b/include/boost/corosio/native/detail/epoll/epoll_traits.hpp index 8b57e9b0d..22d61cc1b 100644 --- a/include/boost/corosio/native/detail/epoll/epoll_traits.hpp +++ b/include/boost/corosio/native/detail/epoll/epoll_traits.hpp @@ -23,6 +23,8 @@ #include #include #include +#include +#include /* epoll backend traits. @@ -92,6 +94,39 @@ struct epoll_traits } }; + // Descriptors are not sockets: sendmsg() fails with ENOTSOCK on a + // pipe or character device, so the write path is writev()/write() + // and SIGPIPE suppression is structurally unavailable -- MSG_NOSIGNAL + // is a send() flag and SO_NOSIGPIPE a socket option. A write to a + // pipe whose read end has closed raises SIGPIPE, exactly as a plain + // write(2) would; callers install SIG_IGN. + struct descriptor_write_policy + { + static ssize_t write(int fd, iovec* iovecs, int count) noexcept + { + ssize_t n; + do + { + n = ::writev(fd, iovecs, count); + } + while (n < 0 && errno == EINTR); + return n; + } + + // Single-buffer fast path: skips the kernel's iov_iter setup. + static ssize_t + write_one(int fd, void const* data, std::size_t size) noexcept + { + ssize_t n; + do + { + n = ::write(fd, data, size); + } + while (n < 0 && errno == EINTR); + return n; + } + }; + struct accept_policy { static int diff --git a/include/boost/corosio/native/detail/epoll/epoll_types.hpp b/include/boost/corosio/native/detail/epoll/epoll_types.hpp index 5ea24afba..ad1623172 100644 --- a/include/boost/corosio/native/detail/epoll/epoll_types.hpp +++ b/include/boost/corosio/native/detail/epoll/epoll_types.hpp @@ -25,6 +25,7 @@ #include #include #include +#include namespace boost::corosio::detail { @@ -41,6 +42,8 @@ class epoll_local_stream_acceptor; class epoll_local_stream_acceptor_service; class epoll_local_datagram_socket; class epoll_local_datagram_service; +class epoll_descriptor; +class epoll_descriptor_service; // --- Stream sockets --- @@ -235,6 +238,29 @@ class epoll_local_stream_acceptor final } }; +// --- Descriptors --- + +class epoll_descriptor final + : public reactor_descriptor< + epoll_descriptor, + epoll_traits, + epoll_descriptor_service, + epoll_tcp_acceptor> +{ + using base_type = reactor_descriptor< + epoll_descriptor, + epoll_traits, + epoll_descriptor_service, + epoll_tcp_acceptor>; + friend epoll_descriptor_service; + +public: + explicit epoll_descriptor(epoll_descriptor_service& svc) noexcept + : base_type(svc) + { + } +}; + // --- Services --- class BOOST_COROSIO_DECL epoll_tcp_service final @@ -351,6 +377,24 @@ class BOOST_COROSIO_DECL epoll_local_stream_acceptor_service final } }; +class BOOST_COROSIO_DECL epoll_descriptor_service final + : public reactor_descriptor_service< + epoll_descriptor_service, + epoll_traits, + epoll_descriptor> +{ + using base_type = reactor_descriptor_service< + epoll_descriptor_service, + epoll_traits, + epoll_descriptor>; + +public: + explicit epoll_descriptor_service(capy::execution_context& ctx) + : base_type(ctx) + { + } +}; + } // namespace boost::corosio::detail #endif // BOOST_COROSIO_HAS_EPOLL diff --git a/include/boost/corosio/native/detail/kqueue/kqueue_traits.hpp b/include/boost/corosio/native/detail/kqueue/kqueue_traits.hpp index cfbb163a4..62114097f 100644 --- a/include/boost/corosio/native/detail/kqueue/kqueue_traits.hpp +++ b/include/boost/corosio/native/detail/kqueue/kqueue_traits.hpp @@ -103,6 +103,39 @@ struct kqueue_traits } }; + // Descriptors are not sockets: sendmsg() fails with ENOTSOCK on a + // pipe or character device, so the write path is writev()/write() + // and SIGPIPE suppression is structurally unavailable -- MSG_NOSIGNAL + // is a send() flag and SO_NOSIGPIPE a socket option. A write to a + // pipe whose read end has closed raises SIGPIPE, exactly as a plain + // write(2) would; callers install SIG_IGN. + struct descriptor_write_policy + { + static ssize_t write(int fd, iovec* iovecs, int count) noexcept + { + ssize_t n; + do + { + n = ::writev(fd, iovecs, count); + } + while (n < 0 && errno == EINTR); + return n; + } + + // Single-buffer fast path: skips the kernel's iov_iter setup. + static ssize_t + write_one(int fd, void const* data, std::size_t size) noexcept + { + ssize_t n; + do + { + n = ::write(fd, data, size); + } + while (n < 0 && errno == EINTR); + return n; + } + }; + struct accept_policy { static int diff --git a/include/boost/corosio/native/detail/kqueue/kqueue_types.hpp b/include/boost/corosio/native/detail/kqueue/kqueue_types.hpp index 071a2f128..3e18accc2 100644 --- a/include/boost/corosio/native/detail/kqueue/kqueue_types.hpp +++ b/include/boost/corosio/native/detail/kqueue/kqueue_types.hpp @@ -25,6 +25,7 @@ #include #include #include +#include namespace boost::corosio::detail { @@ -41,6 +42,8 @@ class kqueue_local_stream_acceptor; class kqueue_local_stream_acceptor_service; class kqueue_local_datagram_socket; class kqueue_local_datagram_service; +class kqueue_descriptor; +class kqueue_descriptor_service; // --- Stream sockets --- @@ -238,6 +241,29 @@ class kqueue_local_stream_acceptor final } }; +// --- Descriptors --- + +class kqueue_descriptor final + : public reactor_descriptor< + kqueue_descriptor, + kqueue_traits, + kqueue_descriptor_service, + kqueue_tcp_acceptor> +{ + using base_type = reactor_descriptor< + kqueue_descriptor, + kqueue_traits, + kqueue_descriptor_service, + kqueue_tcp_acceptor>; + friend kqueue_descriptor_service; + +public: + explicit kqueue_descriptor(kqueue_descriptor_service& svc) noexcept + : base_type(svc) + { + } +}; + // --- Services --- class BOOST_COROSIO_DECL kqueue_tcp_service final @@ -358,6 +384,24 @@ class BOOST_COROSIO_DECL kqueue_local_stream_acceptor_service final } }; +class BOOST_COROSIO_DECL kqueue_descriptor_service final + : public reactor_descriptor_service< + kqueue_descriptor_service, + kqueue_traits, + kqueue_descriptor> +{ + using base_type = reactor_descriptor_service< + kqueue_descriptor_service, + kqueue_traits, + kqueue_descriptor>; + +public: + explicit kqueue_descriptor_service(capy::execution_context& ctx) + : base_type(ctx) + { + } +}; + } // namespace boost::corosio::detail #endif // BOOST_COROSIO_HAS_KQUEUE diff --git a/include/boost/corosio/native/detail/posix/posix_random_access_file.hpp b/include/boost/corosio/native/detail/posix/posix_random_access_file.hpp index c43431f47..169c2ec35 100644 --- a/include/boost/corosio/native/detail/posix/posix_random_access_file.hpp +++ b/include/boost/corosio/native/detail/posix/posix_random_access_file.hpp @@ -25,6 +25,7 @@ #include #include #include +#include #include #include #include @@ -269,6 +270,23 @@ posix_random_access_file::release() inline std::error_code posix_random_access_file::assign(native_handle_type handle) noexcept { + // handle >= 0 guard: an unset impl reports native_handle() == -1, and a + // caller-supplied -1 must fail as a bad fd, not a self-assign. + if (handle >= 0 && handle == fd_) + return std::make_error_code(std::errc::invalid_argument); + + // Validate before touching the held fd: a failed assign must leave + // this object unchanged and the caller still owning handle. + if (auto ec = validate_file_fd(handle)) + return ec; + + // cancel() first: an in-flight read_at/write_at's pool-thread + // completion reads fd_ at execution time, not at post time, so + // without this a pending op silently completes against the newly + // adopted file instead of being cancelled. The service's + // close(handle) / destroy() normally pair cancel()+close_file(); + // assign() bypasses that path and must do the same pairing itself. + cancel(); close_file(); fd_ = handle; return {}; diff --git a/include/boost/corosio/native/detail/posix/posix_stream_file.hpp b/include/boost/corosio/native/detail/posix/posix_stream_file.hpp index 0f0d6c092..9c7d4981e 100644 --- a/include/boost/corosio/native/detail/posix/posix_stream_file.hpp +++ b/include/boost/corosio/native/detail/posix/posix_stream_file.hpp @@ -26,6 +26,7 @@ #include #include #include +#include #include #include #include @@ -328,6 +329,23 @@ posix_stream_file::release() inline std::error_code posix_stream_file::assign(native_handle_type handle) noexcept { + // handle >= 0 guard: an unset impl reports native_handle() == -1, and a + // caller-supplied -1 must fail as a bad fd, not a self-assign. + if (handle >= 0 && handle == fd_) + return std::make_error_code(std::errc::invalid_argument); + + // Validate before touching the held fd: a failed assign must leave + // this object unchanged and the caller still owning handle. + if (auto ec = validate_file_fd(handle)) + return ec; + + // cancel() first: an in-flight read/write's pool-thread completion + // reads fd_/offset_ at execution time, not at post time, so without + // this a pending op silently completes against the newly adopted + // file instead of being cancelled. The service's close(handle) / + // destroy() normally pair cancel()+close_file(); assign() bypasses + // that path and must do the same pairing itself. + cancel(); close_file(); fd_ = handle; offset_ = 0; diff --git a/include/boost/corosio/native/detail/reactor/reactor_descriptor.hpp b/include/boost/corosio/native/detail/reactor/reactor_descriptor.hpp new file mode 100644 index 000000000..76168ee66 --- /dev/null +++ b/include/boost/corosio/native/detail/reactor/reactor_descriptor.hpp @@ -0,0 +1,872 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +#ifndef BOOST_COROSIO_NATIVE_DETAIL_REACTOR_REACTOR_DESCRIPTOR_HPP +#define BOOST_COROSIO_NATIVE_DETAIL_REACTOR_REACTOR_DESCRIPTOR_HPP + +#include + +#if BOOST_COROSIO_POSIX + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include +#include +#include +#include + +#include +#include +#include + +/* Reactor-backed implementation of posix_descriptor. + + Deliberately does not derive from reactor_basic_socket: that base + is rooted in native_socket_base, which overrides local_endpoint(), + set_option() and get_option() on its ImplBase. posix_descriptor has + no socket verbs, so the shared logic (init_and_register, register_op, + the cancel/close/release op sweeps) is carried here against five op + slots instead of eight. + + The one behavior that is genuinely new: O_NONBLOCK is armed lazily, + on the first read_some/write_some and never from assign() or wait(). + The flag lives on the shared open file description, so arming it is + visible to every other holder of that description -- which is why a + wait()-only user must never trigger it. + + PARALLEL COPY: init_and_register, register_op, cancel_single_op and + the cancel/close/release op sweeps here mirror the socket versions in + reactor_basic_socket.hpp (init_and_register, register_op, + cancel_single_op, do_cancel, do_close_socket, do_release_socket). + The two are separate code because that base carries native_socket_base + and its socket verbs; they are not separate protocols. A fix to the + cancel/park protocol -- slot claiming under desc_state_.mutex, the + cached-edge replay in register_op, the impl_ref_ pinning during + teardown -- belongs in both files. +*/ + +namespace boost::corosio::detail { + +// ============================================================ +// Op types +// ============================================================ + +/* Descriptor op family. + + Mirrors reactor_stream_ops.hpp. The Acceptor parameter is a + placeholder: reactor_op is parameterized on both a socket and an + acceptor impl type, and the shared completion helpers name + acceptor_impl_ in a branch that a descriptor op never takes but the + compiler still instantiates. Passing the backend's acceptor type (as + reactor_dgram_socket_impl already does) keeps those helpers shared + rather than duplicated here. + + @tparam Traits Backend traits (epoll_traits, kqueue_traits, ...). + @tparam Descriptor The concrete descriptor type (forward-declared). + @tparam Acceptor Placeholder acceptor type for the op base. +*/ + +template +struct reactor_descriptor_base_op : reactor_op +{ + void operator()() override; + void cancel() noexcept override; +}; + +template +struct reactor_descriptor_read_op final + : reactor_read_op> +{}; + +template +struct reactor_descriptor_write_op final + : reactor_write_op< + reactor_descriptor_base_op, + typename Traits::descriptor_write_policy> +{}; + +template +struct reactor_descriptor_wait_op final + : reactor_wait_op> +{ + void operator()() override; +}; + +// --- Deferred implementations (instantiated when Descriptor is complete) --- + +template +void +reactor_descriptor_base_op::operator()() +{ + complete_io_op(*this); +} + +template +void +reactor_descriptor_base_op::cancel() noexcept +{ + // A descriptor op is only ever started against a descriptor impl, so + // the acceptor arm of the stream op's cancel() has no counterpart. + if (this->socket_impl_) + this->socket_impl_->cancel_single_op(*this); + else + this->request_cancel(); +} + +template +void +reactor_descriptor_wait_op::operator()() +{ + complete_wait_op(*this); +} + +// ============================================================ +// Descriptor implementation +// ============================================================ + +/** CRTP base for reactor-backed posix_descriptor implementations. + + Holds the adopted descriptor, its reactor registration state, and + the five op slots (read, write, and one wait per direction). + + @tparam Derived The named final class (CRTP self). + @tparam Traits Backend traits (epoll_traits, kqueue_traits, ...). + @tparam Service The backend's descriptor service type. + @tparam Acceptor Placeholder acceptor type for the op base. +*/ +template +class reactor_descriptor + : public posix_descriptor::implementation + , public std::enable_shared_from_this + , public intrusive_list::node +{ + using base_op = reactor_descriptor_base_op; + using read_op = reactor_descriptor_read_op; + using write_op = reactor_descriptor_write_op; + using wait_op = reactor_descriptor_wait_op; + +protected: + // NOLINTNEXTLINE(bugprone-crtp-constructor-accessibility) + explicit reactor_descriptor(Service& svc) noexcept : svc_(svc) {} + +public: + ~reactor_descriptor() override = default; + + /// Per-descriptor state for persistent reactor registration. + typename Traits::desc_state_type desc_state_; + + // --- Virtual method overrides --- + + std::coroutine_handle<> read_some( + std::coroutine_handle<> h, + capy::executor_ref ex, + buffer_param param, + std::stop_token token, + std::error_code* ec, + std::size_t* bytes_out) override + { + return do_read_some(h, ex, param, token, ec, bytes_out); + } + + std::coroutine_handle<> write_some( + std::coroutine_handle<> h, + capy::executor_ref ex, + buffer_param param, + std::stop_token token, + std::error_code* ec, + std::size_t* bytes_out) override + { + return do_write_some(h, ex, param, token, ec, bytes_out); + } + + std::coroutine_handle<> wait( + std::coroutine_handle<> h, + capy::executor_ref ex, + wait_type w, + std::stop_token token, + std::error_code* ec) override + { + return do_wait(h, ex, w, token, ec); + } + + native_handle_type native_handle() const noexcept override + { + return fd_; + } + + native_handle_type release_descriptor() noexcept override; + + void cancel() noexcept override + { + do_cancel(); + } + + // --- Service-facing (non-virtual) --- + + /** Adopt the fd, initialize descriptor state, and register it. + + @param fd The descriptor to adopt. + + @return The error if the reactor rejects the descriptor, in + which case the implementation is left closed and the caller + retains ownership of @a fd; otherwise a default constructed + error code. + */ + std::error_code init_and_register(int fd) noexcept; + + /// Close the descriptor and cancel pending operations. + void close_descriptor() noexcept; + + /// Cancel a single pending operation, claiming it from its slot. + template + void cancel_single_op(Op& op) noexcept; + +private: + /** Arm O_NONBLOCK, once, before the first speculative syscall. + + Reports an errno rather than an error_code because that is what + the op result model records; the round trip is lossless here + because fcntl only fails with codes make_err passes through. + */ + int arm_nonblocking() noexcept + { + if (nonblocking_) + return 0; + if (auto ec = ensure_nonblocking(fd_)) + return ec.value(); + nonblocking_ = true; + return 0; + } + + std::coroutine_handle<> do_read_some( + std::coroutine_handle<>, + capy::executor_ref, + buffer_param, + std::stop_token const&, + std::error_code*, + std::size_t*); + + std::coroutine_handle<> do_write_some( + std::coroutine_handle<>, + capy::executor_ref, + buffer_param, + std::stop_token const&, + std::error_code*, + std::size_t*); + + std::coroutine_handle<> do_wait( + std::coroutine_handle<>, + capy::executor_ref, + wait_type, + std::stop_token const&, + std::error_code*); + + void do_cancel() noexcept; + + /// Register an op with the reactor, handling cached edge events. + template + void register_op( + Op& op, + reactor_op_base*& desc_slot, + bool& ready_flag, + bool is_write_direction = false) noexcept; + + /// Apply @a fn to each of the five op slots. + template + void for_each_op(Fn fn) noexcept + { + fn(rd_); + fn(wr_); + fn(wait_rd_); + fn(wait_wr_); + fn(wait_er_); + } + + /** Claim every parked op out of its descriptor_state slot. + + @param claimed Receives the claimed ops; must hold five. + @param teardown Also clear the cached edge flags and, if the + state is queued in the scheduler, pin the impl alive. + @param self Keepalive used by @a teardown. + @return The number of ops claimed. + */ + int claim_parked_ops( + reactor_op_base** claimed, + bool teardown, + std::shared_ptr const& self) noexcept; + + /// Post claimed ops to the scheduler, keeping the impl alive. + void post_claimed_ops( + reactor_op_base** claimed, + int count, + std::shared_ptr const& self) noexcept; + + /// Sweep every op slot, then drop the reactor registration. + void quiesce() noexcept; + + reactor_op_base** op_to_desc_slot(base_op& op) noexcept; + + Service& svc_; + int fd_ = -1; + bool nonblocking_ = false; + + read_op rd_; + write_op wr_; + wait_op wait_rd_; + wait_op wait_wr_; + wait_op wait_er_; +}; + +// ============================================================ +// Registration and teardown +// ============================================================ + +template +std::error_code +reactor_descriptor::init_and_register( + int fd) noexcept +{ + fd_ = fd; + desc_state_.fd = fd; + { + // Every slot this type owns; connect_op is deliberately absent, + // a descriptor has no connect operation to park there. + std::lock_guard lock(desc_state_.mutex); + desc_state_.read_op = nullptr; + desc_state_.write_op = nullptr; + desc_state_.wait_read_op = nullptr; + desc_state_.wait_write_op = nullptr; + desc_state_.wait_error_op = nullptr; + } + if (auto ec = svc_.scheduler().register_descriptor(fd, &desc_state_)) + { + // Undo the partial state so a failed adopt is + // indistinguishable from a closed implementation. + fd_ = -1; + desc_state_.fd = -1; + desc_state_.registered_events = 0; + return ec; + } + return {}; +} + +template +int +reactor_descriptor::claim_parked_ops( + reactor_op_base** claimed, + bool teardown, + std::shared_ptr const& self) noexcept +{ + int count = 0; + std::lock_guard lock(desc_state_.mutex); + for (auto** slot : + {&desc_state_.read_op, &desc_state_.write_op, + &desc_state_.wait_read_op, &desc_state_.wait_write_op, + &desc_state_.wait_error_op}) + { + if (auto* c = std::exchange(*slot, nullptr)) + claimed[count++] = c; + } + if (teardown) + { + desc_state_.read_ready = false; + desc_state_.write_ready = false; + + // Must be set under the same lock that invoke_deferred_io clears + // is_enqueued_ under, or the impl could be destroyed while the + // scheduler still holds the queued descriptor_state. + if (desc_state_.is_enqueued_.load(std::memory_order_acquire)) + desc_state_.impl_ref_ = self; + } + return count; +} + +template +void +reactor_descriptor::post_claimed_ops( + reactor_op_base** claimed, + int count, + std::shared_ptr const& self) noexcept +{ + for (int i = 0; i < count; ++i) + { + claimed[i]->impl_ptr = self; + svc_.post(claimed[i]); + svc_.work_finished(); + } +} + +template +void +reactor_descriptor::do_cancel() noexcept +{ + auto self = this->weak_from_this().lock(); + if (!self) + return; + + for_each_op([](auto& op) { op.request_cancel(); }); + + reactor_op_base* claimed[5]; + int const count = claim_parked_ops(claimed, /*teardown=*/false, self); + post_claimed_ops(claimed, count, self); +} + +template +void +reactor_descriptor::quiesce() noexcept +{ + auto self = this->weak_from_this().lock(); + if (self) + { + for_each_op([](auto& op) { op.request_cancel(); }); + + reactor_op_base* claimed[5]; + int const count = claim_parked_ops(claimed, /*teardown=*/true, self); + post_claimed_ops(claimed, count, self); + } + + if (fd_ >= 0 && desc_state_.registered_events != 0) + svc_.scheduler().deregister_descriptor(fd_); + + desc_state_.registered_events = 0; + // The next adopted fd starts from an unknown flag state. + nonblocking_ = false; +} + +template +void +reactor_descriptor:: + close_descriptor() noexcept +{ + quiesce(); + + if (fd_ >= 0) + { + ::close(fd_); + fd_ = -1; + } + desc_state_.fd = -1; +} + +template +native_handle_type +reactor_descriptor:: + release_descriptor() noexcept +{ + quiesce(); + + // Do NOT close -- the caller takes ownership. + native_handle_type released = fd_; + fd_ = -1; + desc_state_.fd = -1; + return released; +} + +// ============================================================ +// Op registration and per-op cancellation +// ============================================================ + +template +template +void +reactor_descriptor::register_op( + Op& op, + reactor_op_base*& desc_slot, + bool& ready_flag, + bool is_write_direction) noexcept +{ + svc_.work_started(); + + std::lock_guard lock(desc_state_.mutex); + bool io_done = false; + if (ready_flag) + { + ready_flag = false; + op.perform_io(); + io_done = (op.errn != EAGAIN && op.errn != EWOULDBLOCK); + if (!io_done) + op.errn = 0; + } + + if (io_done || op.cancelled.load(std::memory_order_acquire)) + { + svc_.post(&op); + svc_.work_finished(); + } + else + { + desc_slot = &op; + + // Select must rebuild its fd_sets when a write-direction op + // is parked, so select() watches for writability. Compiled + // away to nothing for epoll and kqueue. + if constexpr (Service::needs_write_notification) + { + if (is_write_direction) + svc_.scheduler().notify_reactor(); + } + } +} + +template +reactor_op_base** +reactor_descriptor::op_to_desc_slot( + base_op& op) noexcept +{ + if (&op == static_cast(&rd_)) + return &desc_state_.read_op; + if (&op == static_cast(&wr_)) + return &desc_state_.write_op; + if (&op == static_cast(&wait_rd_)) + return &desc_state_.wait_read_op; + if (&op == static_cast(&wait_wr_)) + return &desc_state_.wait_write_op; + if (&op == static_cast(&wait_er_)) + return &desc_state_.wait_error_op; + return nullptr; +} + +template +template +void +reactor_descriptor::cancel_single_op( + Op& op) noexcept +{ + auto self = this->weak_from_this().lock(); + if (!self) + return; + + op.request_cancel(); + + reactor_op_base** desc_op_ptr = op_to_desc_slot(op); + if (!desc_op_ptr) + return; + + reactor_op_base* claimed = nullptr; + { + std::lock_guard lock(desc_state_.mutex); + if (*desc_op_ptr == &op) + claimed = std::exchange(*desc_op_ptr, nullptr); + // Not in the slot: request_cancel() above already set + // op.cancelled, which register_op consults before parking + // and the completion decode consults on delivery. Latching + // a descriptor flag here instead would outlive this op and + // cancel the next wait in the same direction. + } + if (claimed) + { + op.impl_ptr = self; + svc_.post(&op); + svc_.work_finished(); + } +} + +// ============================================================ +// I/O dispatch +// ============================================================ + +template +std::coroutine_handle<> +reactor_descriptor::do_read_some( + std::coroutine_handle<> h, + capy::executor_ref ex, + buffer_param param, + std::stop_token const& token, + std::error_code* ec, + std::size_t* bytes_out) +{ + auto& op = rd_; + op.reset(); + op.h = h; + op.ex = ex; + op.ec_out = ec; + op.bytes_out = bytes_out; + + // Closed-object contract: complete with bad_file_descriptor without + // touching the kernel or the unregistered descriptor state. + if (fd_ < 0) + { + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + op.complete(EBADF, 0); + svc_.post(&op); + return std::noop_coroutine(); + } + + capy::mutable_buffer bufs[read_op::max_buffers]; + op.iovec_count = + static_cast(param.copy_to(bufs, read_op::max_buffers)); + + if (op.iovec_count == 0 || (op.iovec_count == 1 && bufs[0].size() == 0)) + { + op.empty_buffer_read = true; + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + op.complete(0, 0); + svc_.post(&op); + return std::noop_coroutine(); + } + + // The first transferring operation is what arms O_NONBLOCK; assign() + // and wait() never do. + if (int const nerr = arm_nonblocking()) + { + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + op.complete(nerr, 0); + svc_.post(&op); + return std::noop_coroutine(); + } + + for (int i = 0; i < op.iovec_count; ++i) + { + op.iovecs[i].iov_base = bufs[i].data(); + op.iovecs[i].iov_len = bufs[i].size(); + } + + // Speculative read; the single-buffer case uses read() so the kernel + // skips the readv iov_iter setup. + ssize_t n; + if (op.iovec_count == 1) + { + do + { + n = ::read(fd_, bufs[0].data(), bufs[0].size()); + } + while (n < 0 && errno == EINTR); + } + else + { + do + { + n = ::readv(fd_, op.iovecs, op.iovec_count); + } + while (n < 0 && errno == EINTR); + } + + if (n >= 0 || (errno != EAGAIN && errno != EWOULDBLOCK)) + { + int err = (n < 0) ? errno : 0; + auto bytes = (n > 0) ? static_cast(n) : std::size_t(0); + + if (svc_.scheduler().try_consume_inline_budget()) + { + if (err) + *ec = make_err(err); + else if (n == 0) + *ec = capy::error::eof; + else + *ec = {}; + *bytes_out = bytes; + op.cont.h = h; + return dispatch_coro(ex, op.cont); + } + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + op.complete(err, bytes); + svc_.post(&op); + return std::noop_coroutine(); + } + + // EAGAIN — register with reactor + op.fd = fd_; + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + + register_op(op, desc_state_.read_op, desc_state_.read_ready); + return std::noop_coroutine(); +} + +template +std::coroutine_handle<> +reactor_descriptor::do_write_some( + std::coroutine_handle<> h, + capy::executor_ref ex, + buffer_param param, + std::stop_token const& token, + std::error_code* ec, + std::size_t* bytes_out) +{ + auto& op = wr_; + op.reset(); + op.h = h; + op.ex = ex; + op.ec_out = ec; + op.bytes_out = bytes_out; + + if (fd_ < 0) + { + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + op.complete(EBADF, 0); + svc_.post(&op); + return std::noop_coroutine(); + } + + capy::mutable_buffer bufs[write_op::max_buffers]; + op.iovec_count = + static_cast(param.copy_to(bufs, write_op::max_buffers)); + + if (op.iovec_count == 0 || (op.iovec_count == 1 && bufs[0].size() == 0)) + { + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + op.complete(0, 0); + svc_.post(&op); + return std::noop_coroutine(); + } + + if (int const nerr = arm_nonblocking()) + { + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + op.complete(nerr, 0); + svc_.post(&op); + return std::noop_coroutine(); + } + + for (int i = 0; i < op.iovec_count; ++i) + { + op.iovecs[i].iov_base = bufs[i].data(); + op.iovecs[i].iov_len = bufs[i].size(); + } + + // Speculative write; the single-buffer case skips the iov_iter setup. + ssize_t n; + if (op.iovec_count == 1) + { + n = write_op::write_policy::write_one( + fd_, bufs[0].data(), bufs[0].size()); + } + else + { + n = write_op::write_policy::write(fd_, op.iovecs, op.iovec_count); + } + + if (n >= 0 || (errno != EAGAIN && errno != EWOULDBLOCK)) + { + int err = (n < 0) ? errno : 0; + auto bytes = (n > 0) ? static_cast(n) : std::size_t(0); + + if (svc_.scheduler().try_consume_inline_budget()) + { + *ec = err ? make_err(err) : std::error_code{}; + *bytes_out = bytes; + op.cont.h = h; + return dispatch_coro(ex, op.cont); + } + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + op.complete(err, bytes); + svc_.post(&op); + return std::noop_coroutine(); + } + + // EAGAIN — register with reactor + op.fd = fd_; + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + + register_op(op, desc_state_.write_op, desc_state_.write_ready, true); + return std::noop_coroutine(); +} + +template +std::coroutine_handle<> +reactor_descriptor::do_wait( + std::coroutine_handle<> h, + capy::executor_ref ex, + wait_type w, + std::stop_token const& token, + std::error_code* ec) +{ + // Pick refs up-front to avoid duplicating the register_op call. + wait_op* op_ptr; + reactor_op_base** desc_slot_ptr; + std::uint32_t event; + + if (w == wait_type::read) + { + op_ptr = &wait_rd_; + desc_slot_ptr = &desc_state_.wait_read_op; + event = reactor_event_read; + } + else if (w == wait_type::write) + { + op_ptr = &wait_wr_; + desc_slot_ptr = &desc_state_.wait_write_op; + event = reactor_event_write; + } + else // wait_type::error + { + op_ptr = &wait_er_; + desc_slot_ptr = &desc_state_.wait_error_op; + event = reactor_event_error; + } + + auto& op = *op_ptr; + + // Speculative probe: an edge-triggered reactor cannot report a + // condition that already holds, so a wait initiated on an already + // ready descriptor would otherwise park forever. No syscall here + // modifies the descriptor -- in particular O_NONBLOCK is untouched. + int perr = 0; + if (wait_op::probe(fd_, event, perr)) + { + if (svc_.scheduler().try_consume_inline_budget()) + { + *ec = perr ? make_err(perr) : std::error_code{}; + op.cont.h = h; + return dispatch_coro(ex, op.cont); + } + op.reset(); + op.wait_event = event; + op.h = h; + op.ex = ex; + op.ec_out = ec; + op.fd = fd_; + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + op.complete(perr, 0); + svc_.post(&op); + return std::noop_coroutine(); + } + + op.reset(); + op.wait_event = event; + op.h = h; + op.ex = ex; + op.ec_out = ec; + op.fd = fd_; + op.start(token, static_cast(this)); + op.impl_ptr = this->shared_from_this(); + + // Force register_op's ready path so the wait op re-probes under the + // descriptor mutex before parking. A stale write_ready latched at + // registration would otherwise report a full pipe as writable. + bool force_probe = true; + register_op(op, *desc_slot_ptr, force_probe, event == reactor_event_write); + return std::noop_coroutine(); +} + +} // namespace boost::corosio::detail + +#endif // BOOST_COROSIO_POSIX + +#endif // BOOST_COROSIO_NATIVE_DETAIL_REACTOR_REACTOR_DESCRIPTOR_HPP diff --git a/include/boost/corosio/native/detail/reactor/reactor_descriptor_service.hpp b/include/boost/corosio/native/detail/reactor/reactor_descriptor_service.hpp new file mode 100644 index 000000000..d1ae1a617 --- /dev/null +++ b/include/boost/corosio/native/detail/reactor/reactor_descriptor_service.hpp @@ -0,0 +1,171 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +#ifndef BOOST_COROSIO_NATIVE_DETAIL_REACTOR_REACTOR_DESCRIPTOR_SERVICE_HPP +#define BOOST_COROSIO_NATIVE_DETAIL_REACTOR_REACTOR_DESCRIPTOR_SERVICE_HPP + +#include + +#if BOOST_COROSIO_POSIX + +#include +#include +#include +#include +#include +#include + +#include +#include +#include + +/* Reactor-backed descriptor_service. + + assign_descriptor is the validate-before-mutate core the public + assign() contract rests on, modelled on do_assign_fd in + reactor_service_finals.hpp. + + PARALLEL COPY: construct, destroy, close and shutdown here mirror the + same four members of reactor_socket_service.hpp, which this cannot + reuse because it calls close_socket() by name. A fix to the service + lifecycle -- the construct/destroy bookkeeping under state_->mutex_, + or shutdown's deliberate retention of impl_ptrs_ so impls outlive the + scheduler's drain -- belongs in both files. +*/ + +namespace boost::corosio::detail { + +/** CRTP base for reactor-backed descriptor services. + + @tparam Derived The named final service type (CRTP self). + @tparam Traits Backend traits (epoll_traits, kqueue_traits, ...). + @tparam DescFinal The named final descriptor impl type. +*/ +template +class reactor_descriptor_service : public descriptor_service +{ + using scheduler_type = typename Traits::scheduler_type; + using state_type = reactor_service_state; + + friend Derived; + +protected: + // NOLINTNEXTLINE(bugprone-crtp-constructor-accessibility) + explicit reactor_descriptor_service(capy::execution_context& ctx) + : state_( + std::make_unique( + ctx.template use_service())) + { + } + +public: + /// True when a parked write-direction op must wake the reactor. + static constexpr bool needs_write_notification = + Traits::needs_write_notification; + + ~reactor_descriptor_service() override = default; + + std::error_code assign_descriptor( + posix_descriptor::implementation& impl, native_handle_type fd) override; + + void shutdown() override + { + std::lock_guard lock(state_->mutex_); + + while (auto* impl = state_->impl_list_.pop_front()) + impl->close_descriptor(); + + // Don't clear impl_ptrs_ here: the scheduler shuts down after us + // and drains completed_ops_, so every impl must outlive that. + } + + io_object::implementation* construct() override + { + auto impl = std::make_shared(static_cast(*this)); + auto* raw = impl.get(); + + { + std::lock_guard lock(state_->mutex_); + state_->impl_ptrs_.emplace(raw, std::move(impl)); + state_->impl_list_.push_back(raw); + } + + return raw; + } + + void destroy(io_object::implementation* impl) override + { + auto* typed = static_cast(impl); + typed->close_descriptor(); + std::lock_guard lock(state_->mutex_); + state_->impl_list_.remove(typed); + state_->impl_ptrs_.erase(typed); + } + + void close(io_object::handle& h) override + { + static_cast(h.get())->close_descriptor(); + } + + scheduler_type& scheduler() const noexcept + { + return state_->sched_; + } + + void post(scheduler_op* op) + { + state_->sched_.post(op); + } + + void work_started() noexcept + { + state_->sched_.work_started(); + } + + void work_finished() noexcept + { + state_->sched_.work_finished(); + } + +protected: + std::unique_ptr state_; + +private: + reactor_descriptor_service(reactor_descriptor_service const&) = delete; + reactor_descriptor_service& + operator=(reactor_descriptor_service const&) = delete; +}; + +template +std::error_code +reactor_descriptor_service::assign_descriptor( + posix_descriptor::implementation& impl_base, native_handle_type fd) +{ + auto* impl = static_cast(&impl_base); + + // fd >= 0 guard: an unset impl reports native_handle() == -1, and a + // caller-supplied -1 must fail as a bad fd, not a self-assign. + if (fd >= 0 && fd == impl->native_handle()) + return std::make_error_code(std::errc::invalid_argument); + + // Validate before touching the held descriptor: a failed assign + // must leave the object unchanged and the caller owning the fd. + if (auto ec = validate_descriptor_fd(fd)) + return ec; + + impl->close_descriptor(); + + return impl->init_and_register(fd); +} + +} // namespace boost::corosio::detail + +#endif // BOOST_COROSIO_POSIX + +#endif // BOOST_COROSIO_NATIVE_DETAIL_REACTOR_REACTOR_DESCRIPTOR_SERVICE_HPP diff --git a/include/boost/corosio/native/detail/reactor/reactor_descriptor_state.hpp b/include/boost/corosio/native/detail/reactor/reactor_descriptor_state.hpp index 4dda0eb76..c8a8fddd4 100644 --- a/include/boost/corosio/native/detail/reactor/reactor_descriptor_state.hpp +++ b/include/boost/corosio/native/detail/reactor/reactor_descriptor_state.hpp @@ -1,5 +1,6 @@ // // Copyright (c) 2026 Steve Gerbino +// Copyright (c) 2026 Michael Vandeberg // // Distributed under the Boost Software License, Version 1.0. (See accompanying // file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) @@ -152,7 +153,31 @@ reactor_descriptor_state::invoke_deferred_io() { socklen_t len = sizeof(err); if (::getsockopt(fd, SOL_SOCKET, SO_ERROR, &err, &len) < 0) - err = errno; + { + if (errno == ENOTSOCK) + { + // Non-socket fd (pipe, chardev, ...): no SO_ERROR, so + // let the op's own syscall name the real failure. + // Also force the read/write dispatch below to run: an + // edge-triggered EPOLLERR can arrive alone, and + // without this a parked op never calls perform_io() + // and, the edge being one-shot, never gets another + // chance -- a permanent hang. + // + // Assumes at least one parked op's own syscall makes + // non-EAGAIN progress; if every op re-parks with + // EAGAIN this sticky error is never redelivered and + // they hang. No such case is known for pipes -- a + // future non-socket type that hits one should be + // handled here. + err = 0; + ev |= reactor_event_read | reactor_event_write; + } + else + { + err = errno; + } + } // select raises its exceptional set for out-of-band/urgent // data as well as for genuine faults; on a healthy socket the // probe then reads SO_ERROR == 0. Faulting a pending read or diff --git a/include/boost/corosio/native/detail/select/select_traits.hpp b/include/boost/corosio/native/detail/select/select_traits.hpp index 3b08f6faa..08d5fbd01 100644 --- a/include/boost/corosio/native/detail/select/select_traits.hpp +++ b/include/boost/corosio/native/detail/select/select_traits.hpp @@ -25,6 +25,7 @@ #include #include #include +#include #include /* select backend traits. @@ -111,6 +112,39 @@ struct select_traits } }; + // Descriptors are not sockets: sendmsg() fails with ENOTSOCK on a + // pipe or character device, so the write path is writev()/write() + // and SIGPIPE suppression is structurally unavailable -- MSG_NOSIGNAL + // is a send() flag and SO_NOSIGPIPE a socket option. A write to a + // pipe whose read end has closed raises SIGPIPE, exactly as a plain + // write(2) would; callers install SIG_IGN. + struct descriptor_write_policy + { + static ssize_t write(int fd, iovec* iovecs, int count) noexcept + { + ssize_t n; + do + { + n = ::writev(fd, iovecs, count); + } + while (n < 0 && errno == EINTR); + return n; + } + + // Single-buffer fast path: skips the kernel's iov_iter setup. + static ssize_t + write_one(int fd, void const* data, std::size_t size) noexcept + { + ssize_t n; + do + { + n = ::write(fd, data, size); + } + while (n < 0 && errno == EINTR); + return n; + } + }; + struct accept_policy { static int diff --git a/include/boost/corosio/native/detail/select/select_types.hpp b/include/boost/corosio/native/detail/select/select_types.hpp index cb919a7a4..df21ee80b 100644 --- a/include/boost/corosio/native/detail/select/select_types.hpp +++ b/include/boost/corosio/native/detail/select/select_types.hpp @@ -25,6 +25,7 @@ #include #include #include +#include namespace boost::corosio::detail { @@ -41,6 +42,8 @@ class select_local_stream_acceptor; class select_local_stream_acceptor_service; class select_local_datagram_socket; class select_local_datagram_service; +class select_descriptor; +class select_descriptor_service; // --- Stream sockets --- @@ -238,6 +241,29 @@ class select_local_stream_acceptor final } }; +// --- Descriptors --- + +class select_descriptor final + : public reactor_descriptor< + select_descriptor, + select_traits, + select_descriptor_service, + select_tcp_acceptor> +{ + using base_type = reactor_descriptor< + select_descriptor, + select_traits, + select_descriptor_service, + select_tcp_acceptor>; + friend select_descriptor_service; + +public: + explicit select_descriptor(select_descriptor_service& svc) noexcept + : base_type(svc) + { + } +}; + // --- Services --- class BOOST_COROSIO_DECL select_tcp_service final @@ -358,6 +384,24 @@ class BOOST_COROSIO_DECL select_local_stream_acceptor_service final } }; +class BOOST_COROSIO_DECL select_descriptor_service final + : public reactor_descriptor_service< + select_descriptor_service, + select_traits, + select_descriptor> +{ + using base_type = reactor_descriptor_service< + select_descriptor_service, + select_traits, + select_descriptor>; + +public: + explicit select_descriptor_service(capy::execution_context& ctx) + : base_type(ctx) + { + } +}; + } // namespace boost::corosio::detail #endif // BOOST_COROSIO_HAS_SELECT diff --git a/include/boost/corosio/native/detail/uring/uring_descriptor.hpp b/include/boost/corosio/native/detail/uring/uring_descriptor.hpp new file mode 100644 index 000000000..c7f9167ea --- /dev/null +++ b/include/boost/corosio/native/detail/uring/uring_descriptor.hpp @@ -0,0 +1,569 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +#ifndef BOOST_COROSIO_NATIVE_DETAIL_URING_URING_DESCRIPTOR_HPP +#define BOOST_COROSIO_NATIVE_DETAIL_URING_URING_DESCRIPTOR_HPP + +#include + +#if BOOST_COROSIO_HAS_URING + +#include +#include +#include +#include +#include +#include + +#include +#include +#include +#include +#include +#include + +#include +#include +#include + +/* io_uring-backed implementation of posix_descriptor. + + Three things differ from the reactor backends and from the other + io_uring services: + + Transfers submit READV/WRITEV at offset -1 so the kernel uses (and + advances) the descriptor's own file position. The file services + pass a real offset; a pipe, tty or character device has none. + + O_NONBLOCK is armed lazily, on the first read_some/write_some and + never from assign() or wait(). Here that is a cancellability + requirement, not only the public contract: a transfer the kernel + would have to block on is punted to an io-wq worker, and a request + running on a worker cannot be cancelled. With the flag set the + kernel reports readiness instead, so every operation stays + cancellable on every kernel version. + + That reporting is what the transfer ops' two-phase shape handles. + An O_NONBLOCK descriptor the kernel cannot retry internally + completes with -EAGAIN; the op then re-arms itself as a poll_add on + the same descriptor and re-submits the transfer when the poll says + ready. The handler makes that decision *before* coro_drain_if_shutdown, + which disarms stop_cb: an op going round again keeps its + cancellation wiring, so a stop_token firing between phases still + reaches the kernel. + + The gap between a CQE and its dispatch is the whole difficulty of + that shape, and two epoch counters close it. Nothing of a transfer + waiting in that gap is in the ring, so neither cancel-by-fd nor a + one-shot stop_callback can reach it; cancel() bumps cancel_epoch_ + and every descriptor change bumps desc_epoch_, and an op whose + snapshot no longer matches completes instead of re-arming. The + descriptor epoch is also what makes the staleness check exact + rather than heuristic: a closed fd number the next assign() gets + back would satisfy a bare fd comparison. + + There is no adopt-time registration. assign() validates and takes + the descriptor; a kernel that refuses it says so at the first + operation, as the public docstring promises. +*/ + +namespace boost::corosio::detail { + +class uring_descriptor; + +/** Advance a two-phase transfer op, or report that it is finished. + + A kernel `EAGAIN` becomes a `poll_add` on the same descriptor, and + the poll's completion re-submits the transfer. Every path that + stops instead leaves @a op carrying a result the completion decode + can read as terminal. + + @param op The op whose CQE just arrived. + @return True when a fresh SQE was submitted, in which case the + caller must neither complete nor disarm the op. +*/ +template +bool uring_descriptor_continue(Op& op) noexcept; + +/** Scatter read via `IORING_OP_READV` at the descriptor's own offset. + + @see uring_descriptor_continue for the `polling` phase. +*/ +struct uring_descriptor_read_op final : uring_file_read_op_base +{ + uring_descriptor* desc = nullptr; + /// True while the submitted SQE is the readiness poll, not the read. + bool polling = false; + /// Owner epochs snapshotted at submission; see uring_descriptor_continue. + std::uint32_t cancel_epoch = 0; + std::uint32_t desc_epoch = 0; + + uring_descriptor_read_op() noexcept : uring_file_read_op_base(&do_handler) + { + prep_func = &do_prep; + } + + static void do_prep(uring_op* base, ::io_uring_sqe* sqe) noexcept + { + auto* self = static_cast(base); + if (self->polling) + ::io_uring_prep_poll_add(sqe, self->fd, POLLIN); + else + uring_file_read_op_base::do_prep(base, sqe); + } + + static void do_handler( + void* owner, + scheduler_op* base, + std::uint32_t bytes, + std::uint32_t error) noexcept; +}; + +/// Gather write via `IORING_OP_WRITEV` at the descriptor's own offset. +struct uring_descriptor_write_op final : uring_file_write_op_base +{ + uring_descriptor* desc = nullptr; + /// True while the submitted SQE is the readiness poll, not the write. + bool polling = false; + /// Owner epochs snapshotted at submission; see uring_descriptor_continue. + std::uint32_t cancel_epoch = 0; + std::uint32_t desc_epoch = 0; + + uring_descriptor_write_op() noexcept : uring_file_write_op_base(&do_handler) + { + prep_func = &do_prep; + } + + static void do_prep(uring_op* base, ::io_uring_sqe* sqe) noexcept + { + auto* self = static_cast(base); + if (self->polling) + ::io_uring_prep_poll_add(sqe, self->fd, POLLOUT); + else + uring_file_write_op_base::do_prep(base, sqe); + } + + static void do_handler( + void* owner, + scheduler_op* base, + std::uint32_t bytes, + std::uint32_t error) noexcept; +}; + +/** Native io_uring implementation of @ref posix_descriptor. + + Holds the adopted descriptor and the five embedded op slots: one + transfer per direction, and one wait per direction. + + @par Thread Safety + Distinct objects: Safe.@n + Shared objects: Unsafe. Each slot carries a single pending + operation, so a descriptor must not have two operations of the + same kind in flight. +*/ +class BOOST_COROSIO_DECL uring_descriptor final + : public posix_descriptor::implementation + , public std::enable_shared_from_this +{ + uring_scheduler* sched_ = nullptr; + int fd_ = -1; + bool nonblocking_ = false; + + // Bumped by cancel() and by every descriptor change respectively. + // A transfer op parked between its EAGAIN CQE and its dispatch is + // invisible to the ring, so these are the only thing that can stop + // it re-arming against an intent it no longer belongs to. + std::atomic cancel_epoch_{0}; + std::atomic desc_epoch_{0}; + + uring_descriptor_read_op rd_; + uring_descriptor_write_op wr_; + uring_wait_op wait_rd_; + uring_wait_op wait_wr_; + uring_wait_op wait_er_; + +public: + explicit uring_descriptor(uring_scheduler& sched) noexcept : sched_(&sched) + { + } + + ~uring_descriptor() override + { + close_descriptor(); + } + + // -- io_stream::implementation -- + + std::coroutine_handle<> read_some( + std::coroutine_handle<> h, + capy::executor_ref ex, + buffer_param buffers, + std::stop_token token, + std::error_code* ec, + std::size_t* bytes) override + { + rd_.prepare( + h, ex, ec, bytes, fd_, /*file_offset=*/-1, sched_, + shared_from_this(), buffers, token); + arm_slot(rd_); + sched_->work_started(); + + // Closed-object contract outranks the zero-length no-op. + if (fd_ < 0) + { + rd_.empty_buffer = false; + rd_.res = -EBADF; + push_completed(&rd_); + return std::noop_coroutine(); + } + + if (rd_.empty_buffer || rd_.cancelled.load(std::memory_order_acquire)) + { + push_completed(&rd_); + return std::noop_coroutine(); + } + + // The first transferring operation is what arms O_NONBLOCK; + // assign() and wait() never do. + if (int const nerr = arm_nonblocking()) + { + rd_.res = -nerr; + push_completed(&rd_); + return std::noop_coroutine(); + } + + uring_submit_op(*sched_, &rd_); + return std::noop_coroutine(); + } + + std::coroutine_handle<> write_some( + std::coroutine_handle<> h, + capy::executor_ref ex, + buffer_param buffers, + std::stop_token token, + std::error_code* ec, + std::size_t* bytes) override + { + wr_.prepare( + h, ex, ec, bytes, fd_, /*file_offset=*/-1, sched_, + shared_from_this(), buffers, token); + arm_slot(wr_); + sched_->work_started(); + + if (fd_ < 0) + { + wr_.empty_buffer = false; + wr_.res = -EBADF; + push_completed(&wr_); + return std::noop_coroutine(); + } + + if (wr_.empty_buffer || wr_.cancelled.load(std::memory_order_acquire)) + { + push_completed(&wr_); + return std::noop_coroutine(); + } + + if (int const nerr = arm_nonblocking()) + { + wr_.res = -nerr; + push_completed(&wr_); + return std::noop_coroutine(); + } + + uring_submit_op(*sched_, &wr_); + return std::noop_coroutine(); + } + + // -- posix_descriptor::implementation -- + + std::coroutine_handle<> wait( + std::coroutine_handle<> h, + capy::executor_ref ex, + wait_type w, + std::stop_token token, + std::error_code* ec) override + { + uring_wait_op* op = nullptr; + int poll_flags = 0; + switch (w) + { + case wait_type::read: + op = &wait_rd_; + poll_flags = POLLIN; + break; + case wait_type::write: + op = &wait_wr_; + poll_flags = POLLOUT; + break; + case wait_type::error: + op = &wait_er_; + // POLLERR, POLLHUP and POLLNVAL are reported whether or not + // they are asked for, so the error wait names only POLLPRI. + poll_flags = POLLPRI; + break; + } + + op->prepare( + h, ex, ec, fd_, sched_, shared_from_this(), poll_flags, token); + sched_->work_started(); + + // No arm_nonblocking() on any branch here: a wait must leave a + // descriptor someone else owns exactly as it found it. + if (fd_ < 0) + { + op->res = -EBADF; + push_completed(op); + return std::noop_coroutine(); + } + + if (op->cancelled.load(std::memory_order_acquire)) + { + push_completed(op); + return std::noop_coroutine(); + } + + uring_submit_op(*sched_, op); + return std::noop_coroutine(); + } + + native_handle_type native_handle() const noexcept override + { + return fd_; + } + + native_handle_type release_descriptor() noexcept override + { + // Flush the cancel while the fd is still open so the kernel + // resolves it before the caller can close and recycle the + // number. Do NOT close -- the caller takes ownership. + if (fd_ >= 0) + sched_->cancel_and_flush(fd_); + native_handle_type released = fd_; + fd_ = -1; + nonblocking_ = false; + desc_epoch_.fetch_add(1, std::memory_order_release); + return released; + } + + void cancel() noexcept override + { + // Bump before the SQE: cancel-by-fd reaches only what the ring + // currently holds, and an op waiting for its handler to run + // holds nothing there. The epoch is what that op consults. + cancel_epoch_.fetch_add(1, std::memory_order_release); + if (fd_ >= 0) + sched_->submit_cancel_by_fd(fd_); + } + + /// Epoch bumped by every @ref cancel. + std::uint32_t cancel_epoch() const noexcept + { + return cancel_epoch_.load(std::memory_order_acquire); + } + + /// Epoch bumped by every change of the held descriptor. + std::uint32_t desc_epoch() const noexcept + { + return desc_epoch_.load(std::memory_order_acquire); + } + + // -- Service-facing (non-virtual) -- + + /** Adopt an already-validated descriptor. + + Resets the lazy-nonblocking latch so a freshly adopted fd is + not assumed to carry the flag from whatever this object held + before. + + @param fd The descriptor to adopt. + */ + void set_descriptor(int fd) noexcept + { + fd_ = fd; + nonblocking_ = false; + desc_epoch_.fetch_add(1, std::memory_order_release); + } + + /// Cancel pending operations and close the descriptor. No-op when + /// already closed. + void close_descriptor() noexcept + { + if (fd_ < 0) + return; + // Both kernel entries below can run a queued pipe write as task + // work; with the reader already gone that raises SIGPIPE. + scoped_sigpipe_block no_sigpipe; + sched_->cancel_and_flush(fd_); + ::close(fd_); + fd_ = -1; + nonblocking_ = false; + desc_epoch_.fetch_add(1, std::memory_order_release); + } + +private: + /** Arm O_NONBLOCK, once, before the first transfer. + + Reports an errno rather than an error_code because the op + result model records a negated errno in `res`; the round trip + is lossless because fcntl only fails with codes make_err + passes through. + */ + int arm_nonblocking() noexcept + { + if (nonblocking_) + return 0; + if (auto ec = ensure_nonblocking(fd_)) + return ec.value(); + nonblocking_ = true; + return 0; + } + + /** Bind a transfer slot to this descriptor for a fresh submission. + + The epoch snapshot taken here is what a later re-arm compares + against; see uring_descriptor_continue. + */ + template + void arm_slot(Op& op) noexcept + { + op.desc = this; + op.polling = false; + op.cancel_epoch = cancel_epoch_.load(std::memory_order_acquire); + op.desc_epoch = desc_epoch_.load(std::memory_order_acquire); + } + + /// Queue an already-counted op for the next dispatch cycle. + void push_completed(scheduler_op* op) noexcept + { + uring_scheduler::lock_type lock(sched_->dispatch_mutex()); + sched_->push_completed_locked(op); + } +}; + +// --- Deferred implementations (need uring_descriptor complete) --- + +template +bool +uring_descriptor_continue(Op& op) noexcept +{ + // A poll CQE carries its revents in `res` -- a small positive + // integer the completion decode would otherwise read as a byte + // count and as success. Every path that abandons the op while that + // value is sitting there has to overwrite it with a terminal one. + bool const mid_poll = op.polling && op.res >= 0; + + if (!op.desc) + { + if (mid_poll) + op.res = -EBADF; + return false; + } + + // Two epochs rather than one, because the two reasons to abandon an + // op name different codes to the caller. + if (op.cancelled.load(std::memory_order_acquire) || + op.desc->cancel_epoch() != op.cancel_epoch) + { + if (mid_poll) + op.res = -ECANCELED; + return false; + } + + if (op.desc->desc_epoch() != op.desc_epoch) + { + if (mid_poll) + op.res = -EBADF; + return false; + } + + if (op.polling) + { + // A poll that failed or was cancelled is the operation's answer. + if (op.res < 0) + return false; + op.polling = false; + } + else if (op.res == -EAGAIN || op.res == -EWOULDBLOCK) + { + op.polling = true; + } + else + { + return false; + } + + // do_one spends a work_finished() on every op it dispatches, so an + // op going round again has to be counted again. + op.sched_->work_started(); + uring_submit_op(*op.sched_, &op); + + // stop_cb is one-shot and has already fired for the SQE that just + // completed, and cancel-by-fd found nothing while this op was out + // of the ring: a cancel racing the submission above would reach no + // kernel request at all, so re-check and drive it here. + if (op.cancelled.load(std::memory_order_acquire) || + op.desc->cancel_epoch() != op.cancel_epoch) + op.sched_->submit_cancel_by_user_data(&op); + return true; +} + +inline void +uring_descriptor_read_op::do_handler( + void* owner, + scheduler_op* base, + std::uint32_t /*bytes*/, + std::uint32_t /*error*/) noexcept +{ + auto* self = static_cast(base); + if (owner != nullptr && uring_descriptor_continue(*self)) + return; + + if (coro_drain_if_shutdown(owner, self)) + return; + + if (self->sched_) + self->sched_->reset_inline_budget(); + + uring_set_result(self, /*is_read=*/true, self->empty_buffer); + if (self->bytes_out) + *self->bytes_out = + self->res >= 0 ? static_cast(self->res) : 0u; + coro_resume(self); +} + +inline void +uring_descriptor_write_op::do_handler( + void* owner, + scheduler_op* base, + std::uint32_t /*bytes*/, + std::uint32_t /*error*/) noexcept +{ + auto* self = static_cast(base); + if (owner != nullptr && uring_descriptor_continue(*self)) + return; + + if (coro_drain_if_shutdown(owner, self)) + return; + + if (self->sched_) + self->sched_->reset_inline_budget(); + + uring_set_result(self, /*is_read=*/false, self->empty_buffer); + if (self->bytes_out) + *self->bytes_out = + self->res >= 0 ? static_cast(self->res) : 0u; + coro_resume(self); +} + +} // namespace boost::corosio::detail + +#endif // BOOST_COROSIO_HAS_URING + +#endif // BOOST_COROSIO_NATIVE_DETAIL_URING_URING_DESCRIPTOR_HPP diff --git a/include/boost/corosio/native/detail/uring/uring_descriptor_service.hpp b/include/boost/corosio/native/detail/uring/uring_descriptor_service.hpp new file mode 100644 index 000000000..712d948ac --- /dev/null +++ b/include/boost/corosio/native/detail/uring/uring_descriptor_service.hpp @@ -0,0 +1,143 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +#ifndef BOOST_COROSIO_NATIVE_DETAIL_URING_URING_DESCRIPTOR_SERVICE_HPP +#define BOOST_COROSIO_NATIVE_DETAIL_URING_URING_DESCRIPTOR_SERVICE_HPP + +#include + +#if BOOST_COROSIO_HAS_URING + +#include +#include +#include +#include +#include +#include + +#include +#include +#include +#include +#include + +/* io_uring-backed descriptor_service. + + assign_descriptor is the validate-before-mutate core the public + assign() contract rests on. It is the reactor version minus the + registration step: io_uring has no adopt-time registration syscall, + so adoption cannot fail once validation has passed, and a kernel + refusal surfaces at the first operation instead. + + Neither uring_socket_service_base nor uring_file_service_base fits: + the first constructs impls with a (service&, scheduler&) ctor and + only cancels on shutdown, the second names its teardown hook + close_file(). The lifecycle here is small enough to carry directly. + + PARALLEL COPY: construct, destroy, close and shutdown here mirror the + same four members of reactor_descriptor_service.hpp, which this + cannot reuse because that template is rooted in the reactor's + scheduler and service state. A fix to the descriptor service + lifecycle -- the construct/destroy bookkeeping, or shutdown's + deliberate retention of the impl map so impls outlive the + scheduler's drain -- belongs in both files. +*/ + +namespace boost::corosio::detail { + +/// Native io_uring descriptor service. Owns every @ref uring_descriptor +/// the context creates. +class BOOST_COROSIO_DECL uring_descriptor_service final + : public descriptor_service +{ + uring_scheduler* sched_ = nullptr; + std::mutex mutex_; + std::unordered_map> + impls_; + +public: + explicit uring_descriptor_service(capy::execution_context& ctx) + : sched_(&ctx.use_service()) + { + } + + std::error_code assign_descriptor( + posix_descriptor::implementation& impl_base, + native_handle_type fd) override + { + auto* impl = static_cast(&impl_base); + + // fd >= 0 guard: an unset impl reports native_handle() == -1, and + // a caller-supplied -1 must fail as a bad fd, not a self-assign. + if (fd >= 0 && fd == impl->native_handle()) + return std::make_error_code(std::errc::invalid_argument); + + // Validate before touching the held descriptor: a failed assign + // must leave the object unchanged and the caller owning the fd. + if (auto ec = validate_descriptor_fd(fd)) + return ec; + + impl->close_descriptor(); + impl->set_descriptor(fd); + return {}; + } + + io_object::implementation* construct() override + { + auto p = std::make_shared(*sched_); + auto* raw = p.get(); + std::lock_guard lock(mutex_); + impls_.emplace(raw, std::move(p)); + return raw; + } + + void destroy(io_object::implementation* p) override + { + if (!p) + return; + auto* impl = static_cast(p); + impl->close_descriptor(); + std::lock_guard lock(mutex_); + impls_.erase(impl); + } + + void close(io_object::handle& h) override + { + if (auto* impl = static_cast(h.get())) + impl->close_descriptor(); + } + + void shutdown() override + { + // Snapshot, then close without the lock held. impls_ is + // deliberately not cleared: the scheduler shuts down after this + // service and drains its completed ops, so every impl must + // outlive that drain. + std::vector> live; + { + std::lock_guard lock(mutex_); + live.reserve(impls_.size()); + for (auto& [raw, p] : impls_) + live.push_back(p); + } + for (auto& p : live) + p->close_descriptor(); + } + +private: + uring_descriptor_service(uring_descriptor_service const&) = delete; + uring_descriptor_service& + operator=(uring_descriptor_service const&) = delete; +}; + +} // namespace boost::corosio::detail + +#endif // BOOST_COROSIO_HAS_URING + +#endif // BOOST_COROSIO_NATIVE_DETAIL_URING_URING_DESCRIPTOR_SERVICE_HPP diff --git a/include/boost/corosio/native/detail/uring/uring_random_access_file.hpp b/include/boost/corosio/native/detail/uring/uring_random_access_file.hpp index 773b049bf..361bc123c 100644 --- a/include/boost/corosio/native/detail/uring/uring_random_access_file.hpp +++ b/include/boost/corosio/native/detail/uring/uring_random_access_file.hpp @@ -1,5 +1,6 @@ // // Copyright (c) 2026 Steve Gerbino +// Copyright (c) 2026 Michael Vandeberg // // Distributed under the Boost Software License, Version 1.0. (See accompanying // file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) @@ -20,6 +21,7 @@ #include #include #include +#include #include #include @@ -154,6 +156,20 @@ class BOOST_COROSIO_DECL uring_random_access_file final std::error_code assign(native_handle_type handle) noexcept override { + // handle >= 0 guard: an unset impl reports native_handle() == -1, + // and a caller-supplied -1 must fail as a bad fd, not a + // self-assign. + if (handle >= 0 && handle == fd_) + return std::make_error_code(std::errc::invalid_argument); + + // Validate before touching the held fd: a failed assign must + // leave this object unchanged and the caller still owning handle. + if (auto ec = validate_file_fd(handle)) + return ec; + + // No cancel() before this, unlike the POSIX twin: close_file() + // calls cancel_and_flush(fd_) itself, so the mandatory + // cancel-before-close pairing is already there. close_file(); fd_ = handle; return {}; diff --git a/include/boost/corosio/native/detail/uring/uring_socket_ops.hpp b/include/boost/corosio/native/detail/uring/uring_socket_ops.hpp index 0b564aa7d..bca726b68 100644 --- a/include/boost/corosio/native/detail/uring/uring_socket_ops.hpp +++ b/include/boost/corosio/native/detail/uring/uring_socket_ops.hpp @@ -636,17 +636,39 @@ struct uring_wait_op : uring_op // way — SO_ERROR, or EIO when the kernel has none — instead of // completing wait(error) with an empty, benign-looking code. // OOB (POLLPRI) is a readiness signal, not an error. + // + // POLLHUP is a fault only for wait(error). A readiness wait -- + // the POLLIN/POLLOUT flag sets -- treats it as ready: a pipe or + // fully shut-down socket whose peer hung up is readable-at-EOF, + // and the read or write that follows names the condition. The + // reactors' poll() probe already reports it that way, so + // without this io_uring alone answers the wait-then-read idiom + // with a spurious EIO. + bool const readiness_wait = + (self->poll_flags & (POLLIN | POLLOUT)) != 0; + int const fault_bits = readiness_wait + ? (POLLERR | POLLNVAL) + : (POLLERR | POLLHUP | POLLNVAL); + std::error_code ec{}; if (self->res < 0) { ec = make_err(-self->res); } - else if (self->res & (POLLERR | POLLHUP | POLLNVAL)) + else if (self->res & fault_bits) { int so_err = 0; socklen_t len = sizeof(so_err); if (::getsockopt(self->fd, SOL_SOCKET, SO_ERROR, &so_err, &len) < 0) - so_err = errno; + { + // A non-socket (pipe, chardev, ...) has no SO_ERROR and + // fails the probe with ENOTSOCK; reporting that would + // name the probe rather than the fault, so fall through + // to the EIO substitution below. Every other failure + // (EBADF from a concurrent close, say) still reports + // itself, unchanged on the socket hot path. + so_err = (errno == ENOTSOCK) ? 0 : errno; + } if (so_err == 0) so_err = EIO; ec = make_err(so_err); diff --git a/include/boost/corosio/native/detail/uring/uring_stream_file.hpp b/include/boost/corosio/native/detail/uring/uring_stream_file.hpp index 05bb2ccf6..99f8509c1 100644 --- a/include/boost/corosio/native/detail/uring/uring_stream_file.hpp +++ b/include/boost/corosio/native/detail/uring/uring_stream_file.hpp @@ -1,5 +1,6 @@ // // Copyright (c) 2026 Steve Gerbino +// Copyright (c) 2026 Michael Vandeberg // // Distributed under the Boost Software License, Version 1.0. (See accompanying // file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) @@ -20,6 +21,7 @@ #include #include #include +#include #include #include @@ -161,6 +163,20 @@ class BOOST_COROSIO_DECL uring_stream_file final std::error_code assign(native_handle_type handle) noexcept override { + // handle >= 0 guard: an unset impl reports native_handle() == -1, + // and a caller-supplied -1 must fail as a bad fd, not a + // self-assign. + if (handle >= 0 && handle == fd_) + return std::make_error_code(std::errc::invalid_argument); + + // Validate before touching the held fd: a failed assign must + // leave this object unchanged and the caller still owning handle. + if (auto ec = validate_file_fd(handle)) + return ec; + + // No cancel() before this, unlike the POSIX twin: close_file() + // calls cancel_and_flush(fd_) itself, so the mandatory + // cancel-before-close pairing is already there. close_file(); fd_ = handle; return {}; diff --git a/include/boost/corosio/native/detail/validate_fd.hpp b/include/boost/corosio/native/detail/validate_fd.hpp index 39c4e0c61..fac25449b 100644 --- a/include/boost/corosio/native/detail/validate_fd.hpp +++ b/include/boost/corosio/native/detail/validate_fd.hpp @@ -1,5 +1,6 @@ // // Copyright (c) 2026 Steve Gerbino +// Copyright (c) 2026 Michael Vandeberg // // Distributed under the Boost Software License, Version 1.0. (See accompanying // file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) @@ -19,7 +20,9 @@ #include #include +#include #include +#include namespace boost::corosio::detail { @@ -65,6 +68,114 @@ validate_socket_fd(int fd, int expected_type, bool is_ip) noexcept return {}; } +/** Validate a caller-supplied fd for adoption by @ref posix_descriptor. + + Non-mutating: interrogates the fd without changing any of its + flags, so a rejected fd goes back to the caller untouched. In + particular `O_NONBLOCK` is not applied here — see + @ref ensure_nonblocking. + + The file-type test is a reject-list, not an accept-list. The + flagship descriptor kinds -- eventfd, timerfd, inotify, pidfd -- + are anonymous inodes whose `st_mode` type bits are all zero, so + an accept-list would silently reject exactly the fds this type + exists to carry. + + @param fd The descriptor to validate. + @return Empty on success; `EBADF` for a closed or negative fd, + `operation_not_supported` for a regular file, directory or + block device, or the `errno` reported by `fstat`. +*/ +inline std::error_code +validate_descriptor_fd(int fd) noexcept +{ + if (fd < 0) + return make_err(EBADF); + + struct stat st{}; + if (::fstat(fd, &st) != 0) + return make_err(errno); + + // Regular files, block devices and directories are the province of + // stream_file / random_access_file, whose assign() already adopts + // them; a reactor cannot report readiness for them anyway. + switch (st.st_mode & S_IFMT) + { + case S_IFREG: + case S_IFBLK: + case S_IFDIR: + return std::make_error_code(std::errc::operation_not_supported); + default: + return {}; + } +} + +/** Validate a caller-supplied fd for adoption by a file object. + + Non-mutating. Accepts the kinds a file object can position and + read: regular files, block devices, and character devices such + as /dev/null and /dev/zero. + + This is an accept-list, the inverse of @ref validate_descriptor_fd's + reject-list: a file object needs a positionable fd, and the + anonymous inodes that motivate the descriptor reject-list are + exactly what a file object cannot use. + + `S_IFCHR` is deliberately broad: it admits `/dev/null` and + `/dev/zero`, but also non-seekable character devices such as a + tty. Those pass here and then fail loudly at first I/O on the + POSIX backends, where `preadv`/`pwritev` report `ESPIPE`. + + @param fd The descriptor to validate. + @return Empty on success; `EBADF` for a closed or negative fd, + `operation_not_supported` for a directory or a descriptor + with no file position, or the `errno` from `fstat`. +*/ +inline std::error_code +validate_file_fd(int fd) noexcept +{ + if (fd < 0) + return make_err(EBADF); + + struct stat st{}; + if (::fstat(fd, &st) != 0) + return make_err(errno); + + switch (st.st_mode & S_IFMT) + { + case S_IFREG: + case S_IFBLK: + case S_IFCHR: + return {}; + default: + return std::make_error_code(std::errc::operation_not_supported); + } +} + +/** Put a descriptor into non-blocking mode, idempotently. + + Called lazily on the first `read_some` / `write_some`, never from + `assign()`. The change is permanent: `O_NONBLOCK` lives on the + shared open file description, so restoring it later would race + every other holder of that description. A `wait()`-only user + never reaches this function and their fd is never modified. + + @param fd The descriptor to modify. + @return Empty on success, otherwise the `errno` from `fcntl`. +*/ +inline std::error_code +ensure_nonblocking(int fd) noexcept +{ + int flags = ::fcntl(fd, F_GETFL, 0); + if (flags < 0) + return make_err(errno); + if (flags & O_NONBLOCK) + return {}; + if (::fcntl(fd, F_SETFL, flags | O_NONBLOCK) < 0) + return make_err(errno); + return {}; +} + } // namespace boost::corosio::detail #endif // BOOST_COROSIO_POSIX diff --git a/include/boost/corosio/native/native.hpp b/include/boost/corosio/native/native.hpp index affe8fa62..ef47ee298 100644 --- a/include/boost/corosio/native/native.hpp +++ b/include/boost/corosio/native/native.hpp @@ -1,5 +1,6 @@ // // Copyright (c) 2026 Steve Gerbino +// Copyright (c) 2026 Michael Vandeberg // // Distributed under the Boost Software License, Version 1.0. (See accompanying // file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) @@ -20,6 +21,7 @@ #include #include #include +#include #include #include #include diff --git a/include/boost/corosio/native/native_posix_descriptor.hpp b/include/boost/corosio/native/native_posix_descriptor.hpp new file mode 100644 index 000000000..5275c205e --- /dev/null +++ b/include/boost/corosio/native/native_posix_descriptor.hpp @@ -0,0 +1,229 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +#ifndef BOOST_COROSIO_NATIVE_NATIVE_POSIX_DESCRIPTOR_HPP +#define BOOST_COROSIO_NATIVE_NATIVE_POSIX_DESCRIPTOR_HPP + +#include +#include +#include + +#if BOOST_COROSIO_POSIX || defined(BOOST_COROSIO_MRDOCS) + +#ifndef BOOST_COROSIO_MRDOCS +#if BOOST_COROSIO_HAS_EPOLL +#include +#endif + +#if BOOST_COROSIO_HAS_SELECT +#include +#endif + +#if BOOST_COROSIO_HAS_KQUEUE +#include +#endif + +#if BOOST_COROSIO_HAS_URING +// uring_types.hpp does not declare the descriptor: it was added +// after that header, in its own pair of files. +#include +#endif +#endif // !BOOST_COROSIO_MRDOCS + +namespace boost::corosio { + +/** Drives an already-open POSIX descriptor, calling the backend directly. + + This class template inherits from @ref posix_descriptor and + shadows the async operations (`read_some`, `write_some`, `wait`) + with versions that call the backend implementation directly. + This lets the compiler inline through the entire call chain. + + Non-async operations (`assign`, `release`, `close`, `cancel`) + remain unchanged and dispatch through the compiled library. + + A `native_posix_descriptor` IS-A `posix_descriptor` and can be + passed to any function expecting `posix_descriptor&` or + `io_stream&`, in which case virtual dispatch is used + transparently. + + @tparam Backend A backend tag value (e.g., `epoll`) whose type + provides the concrete implementation types. + + @par Thread Safety + Same as @ref posix_descriptor. + + @par Example + @par !example assign_and_wait + + @see posix_descriptor, epoll_t, kqueue_t +*/ +template +class native_posix_descriptor : public posix_descriptor +{ + using backend_type = decltype(Backend); + using impl_type = typename backend_type::descriptor_type; + using service_type = typename backend_type::descriptor_service_type; + + impl_type& get_impl() noexcept + { + return *static_cast(h_.get()); + } + + template + struct native_read_awaitable + : detail::bytes_op_base> + { + native_posix_descriptor& self_; + MutableBufferSequence buffers_; + + native_read_awaitable( + native_posix_descriptor& self, + MutableBufferSequence buffers) noexcept + : self_(self) + , buffers_(std::move(buffers)) + { + } + + std::coroutine_handle<> + dispatch(std::coroutine_handle<> h, capy::executor_ref ex) const + { + return self_.get_impl().read_some( + h, ex, buffers_, this->token_, &this->ec_, &this->bytes_); + } + }; + + template + struct native_write_awaitable + : detail::bytes_op_base> + { + native_posix_descriptor& self_; + ConstBufferSequence buffers_; + + native_write_awaitable( + native_posix_descriptor& self, ConstBufferSequence buffers) noexcept + : self_(self) + , buffers_(std::move(buffers)) + { + } + + std::coroutine_handle<> + dispatch(std::coroutine_handle<> h, capy::executor_ref ex) const + { + return self_.get_impl().write_some( + h, ex, buffers_, this->token_, &this->ec_, &this->bytes_); + } + }; + + struct native_wait_awaitable : detail::void_op_base + { + native_posix_descriptor& self_; + wait_type w_; + + native_wait_awaitable( + native_posix_descriptor& self, wait_type w) noexcept + : self_(self) + , w_(w) + { + } + + std::coroutine_handle<> + dispatch(std::coroutine_handle<> h, capy::executor_ref ex) const + { + return self_.get_impl().wait(h, ex, w_, this->token_, &this->ec_); + } + }; + +public: + /** Construct a native descriptor from an execution context. + + @param ctx The execution context that owns this object. + */ + explicit native_posix_descriptor(capy::execution_context& ctx) + : io_object(handle(ctx, ctx.use_service())) + { + } + + /** Construct a native descriptor from an executor. + + @param ex The executor whose context owns this object. + */ + template + requires(!std::same_as< + std::remove_cvref_t, + native_posix_descriptor>) && + capy::Executor + explicit native_posix_descriptor(Ex const& ex) + : native_posix_descriptor(ex.context()) + { + } + + /// Move construct. + native_posix_descriptor(native_posix_descriptor&&) noexcept = default; + + /// Move assign. + native_posix_descriptor& + operator=(native_posix_descriptor&&) noexcept = default; + + /// Copy construction is disabled; the handle is uniquely owned. + native_posix_descriptor(native_posix_descriptor const&) = delete; + /// Copy assignment is disabled; the handle is uniquely owned. + native_posix_descriptor& + operator=(native_posix_descriptor const&) = delete; + + /** Asynchronously read data from the descriptor. + + Calls the backend implementation directly, bypassing virtual + dispatch. Otherwise identical to @ref io_stream::read_some. + + @param buffers The buffer sequence to read into. + + @return An awaitable yielding `(error_code, std::size_t)`. + */ + template + [[nodiscard]] auto read_some(MB const& buffers) + { + return native_read_awaitable(*this, buffers); + } + + /** Asynchronously write data to the descriptor. + + Calls the backend implementation directly, bypassing virtual + dispatch. Otherwise identical to @ref io_stream::write_some. + + @param buffers The buffer sequence to write from. + + @return An awaitable yielding `(error_code, std::size_t)`. + */ + template + [[nodiscard]] auto write_some(CB const& buffers) + { + return native_write_awaitable(*this, buffers); + } + + /** Wait for readiness without transferring bytes. + + Calls the backend implementation directly, bypassing virtual + dispatch. Otherwise identical to @ref posix_descriptor::wait. + + @param w The direction to wait on. + + @return An awaitable yielding `io_result<>`. + */ + [[nodiscard]] auto wait(wait_type w) + { + return native_wait_awaitable(*this, w); + } +}; + +} // namespace boost::corosio + +#endif // BOOST_COROSIO_POSIX || BOOST_COROSIO_MRDOCS + +#endif // BOOST_COROSIO_NATIVE_NATIVE_POSIX_DESCRIPTOR_HPP diff --git a/include/boost/corosio/posix_descriptor.hpp b/include/boost/corosio/posix_descriptor.hpp new file mode 100644 index 000000000..bc2f031b6 --- /dev/null +++ b/include/boost/corosio/posix_descriptor.hpp @@ -0,0 +1,362 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +#ifndef BOOST_COROSIO_POSIX_DESCRIPTOR_HPP +#define BOOST_COROSIO_POSIX_DESCRIPTOR_HPP + +#include +#include + +#if BOOST_COROSIO_POSIX || defined(BOOST_COROSIO_MRDOCS) + +#include +#include +#include +#include +#include +#include +#include +#include + +#include +#include +#include +#include +#include + +/* Adoption of an already-open pollable POSIX descriptor. + + The two contract points that are not obvious from the + declarations: + + assign() validates before it mutates. A fd rejected by validation + leaves the object holding whatever it held before, pending + operations included, and leaves ownership of the fd with the + caller. A kernel registration refusal is the one exception. The + previous descriptor is already closed by then, so the object is + left closed. + + O_NONBLOCK is applied lazily, at the first read_some/write_some, + and never restored. A wait()-only user never triggers it, which + is what makes adopting STDIN_FILENO safe: flipping the flag would + change the parent shell's terminal, because the flag lives on the + shared open file description, not on the descriptor. +*/ + +namespace boost::corosio { + +/** Drives an already-open POSIX descriptor from an `io_context`. + + Wraps an already-open pollable file descriptor and drives it + from the `io_context`. The kinds in scope are character devices, + `inotify`, `eventfd`, `timerfd`, `pidfd`, pipes, ttys, and socket + kinds corosio does not otherwise wrap. The descriptor must come + from the caller; this type never creates one. + + The type name is deliberately platform-qualified. Portability + comes from the interfaces it implements, not from the name. A + `posix_descriptor` is an @ref io_stream. `capy::read`, + `capy::write`, other `capy::Stream`-constrained algorithms and + TLS layering therefore work on it exactly as they do on a + socket. + + @par Ownership + `assign()` takes ownership and `close()` closes the + descriptor. To integrate with a library that owns the fd, adopt + a `dup()` of it: readiness lives on the open file description, + which both descriptors share. + + @par Descriptor Flags + `assign()` and `wait()` never modify the descriptor. The first + `read_some()` or `write_some()` sets `O_NONBLOCK` and never + restores it. The flag lives on the shared open file + description, so restoring it would race every other holder. A + `dup()` is no escape: the duplicate shares that same description, + so the flag change reaches the other holder anyway. When another + party owns the descriptor and cannot tolerate `O_NONBLOCK`, use + `wait()` -- which never modifies the descriptor -- and do the I/O + yourself. + + @par Rejected Descriptors + Regular files, block devices, and directories are rejected with + `errc::operation_not_supported`. @ref stream_file and + @ref random_access_file adopt regular files and block devices. A + directory is adoptable by no corosio type. Where a kernel refusal + surfaces depends on the backend. The epoll, kqueue, and select + backends register the descriptor during `assign()`, so a refusal + fails there. On select that is `EMFILE` for `fd >= FD_SETSIZE`. + The io_uring backend has no adopt-time registration, so + `assign()` succeeds and takes ownership, and the refusal appears + at the first `read_some()` or `write_some()`. An `assign()`-time + refusal is the one failure that does not preserve the previously + held descriptor. The previous descriptor is already closed by + then, so the object is left closed. + + @par Signals + Writing to a descriptor whose peer has closed raises `SIGPIPE` + in the default disposition -- unlike the socket types, which + suppress it. `MSG_NOSIGNAL` is a `send()` flag with no `writev` + equivalent, and `SO_NOSIGPIPE` is a socket option, so neither + applies to an arbitrary descriptor. Callers must install + `SIG_IGN` for `SIGPIPE` if that is not already the process's + disposition. + + @par Thread Safety + Distinct objects: Safe.@n + Shared objects: Unsafe. A descriptor must not have concurrent + operations of the same type (e.g. two simultaneous reads). One + read and one write may be in flight simultaneously. + + @see io_stream, stream_file, wait_type +*/ +class BOOST_COROSIO_DECL posix_descriptor : public io_stream +{ +public: + /** Define backend hooks for descriptor operations. + + Platform backends (epoll, kqueue, select, io_uring) derive + from this to implement descriptor I/O. + */ + struct implementation : io_stream::implementation + { + /** Initiate an asynchronous wait for descriptor readiness. + + Completes when the descriptor becomes ready in the + given direction, or an error condition is reported. No + bytes are transferred and no descriptor flag is changed. + + @param h Coroutine handle to resume on completion. + @param ex Executor for dispatching the completion. + @param w The direction to wait on. + @param token Stop token for cancellation. + @param ec Output error code. + @return Coroutine handle to resume immediately. + */ + virtual std::coroutine_handle<> wait( + std::coroutine_handle<> h, + capy::executor_ref ex, + wait_type w, + std::stop_token token, + std::error_code* ec) = 0; + + /// Return the platform descriptor, or -1 when not open. + virtual native_handle_type native_handle() const noexcept = 0; + + /** Release ownership of the native descriptor. + + Stops tracking the descriptor and cancels its pending + operations, without closing it. The caller takes + ownership. + + @return The native descriptor. + */ + virtual native_handle_type release_descriptor() noexcept = 0; + + /** Request cancellation of pending asynchronous operations. + + All outstanding operations complete with a code that + compares equal to `capy::cond::canceled`. + */ + virtual void cancel() noexcept = 0; + }; + + /// Represent the awaitable returned by @ref wait. + struct wait_awaitable : detail::void_op_base + { + private: + friend posix_descriptor; + + wait_awaitable(posix_descriptor& d, wait_type w) noexcept : d_(d), w_(w) + { + } + + friend detail::void_op_base; + + posix_descriptor& d_; + wait_type w_; + + std::coroutine_handle<> + dispatch(std::coroutine_handle<> h, capy::executor_ref ex) const + { + return d_.get().wait(h, ex, w_, token_, &ec_); + } + }; + + /** Destructor. + + Closes the descriptor if open, cancelling pending operations. + */ + ~posix_descriptor() override; + + /** Construct from an execution context. + + @param ctx The execution context that owns this object. + */ + explicit posix_descriptor(capy::execution_context& ctx); + + /** Construct from an executor. + + The overload excludes `posix_descriptor` itself so that it + cannot displace the move constructor. + + @tparam Ex A type satisfying `capy::Executor`. + @param ex The executor whose context owns this object. + */ + template + requires(!std::same_as, posix_descriptor>) && + capy::Executor + explicit posix_descriptor(Ex const& ex) : posix_descriptor(ex.context()) + { + } + + /** Move constructor. + + @param other The object to move from. + @pre No awaitables returned by @p other's methods exist. + */ + posix_descriptor(posix_descriptor&& other) noexcept + : io_object(std::move(other)) + { + } + + /** Move assignment. + + @param other The object to move from. + @return `*this`. + @pre No awaitables returned by either object's methods exist. + */ + posix_descriptor& operator=(posix_descriptor&& other) noexcept + { + io_object::operator=(std::move(other)); + return *this; + } + + /// Copy construction is disabled; the descriptor is uniquely owned. + posix_descriptor(posix_descriptor const&) = delete; + /// Copy assignment is disabled; the descriptor is uniquely owned. + posix_descriptor& operator=(posix_descriptor const&) = delete; + + /** Adopt an existing native descriptor. + + Validation runs before anything is mutated or closed. When + validation rejects @p fd the object still holds whatever + descriptor and pending operations it held before, and the + caller still owns @p fd. On success the object takes + ownership and @p fd is closed by `close()` or the destructor. + + No descriptor flag is modified here, `O_NONBLOCK` included. + + @param fd The native descriptor to adopt. + + @return `errc::invalid_argument` when @p fd is the + descriptor this object already holds. + `errc::bad_file_descriptor` when @p fd is negative or + closed. `errc::operation_not_supported` when @p fd names + a regular file, block device, or directory. Otherwise the + `errno` reported by the kernel, or an empty code. + + @par Exception Safety + Throws nothing. The strong guarantee covers validation + failure only. A kernel registration refusal can occur only + after validation passes, and only on the backends that + register at adopt time (epoll, kqueue, select). The previous + descriptor is already closed by then, so the object is left + closed and @p fd stays with the caller. + + @see release + */ + [[nodiscard]] std::error_code assign(native_handle_type fd) noexcept; + + /** Release ownership of the native descriptor. + + The object becomes not-open and pending operations are + cancelled. The caller is responsible for closing the result. + + @return The native descriptor. + + @throws std::system_error `errc::bad_file_descriptor` if the + object is not open. + + @post `is_open() == false` + */ + native_handle_type release(); + + /** Close the descriptor. + + Pending operations complete with a code that compares equal + to `capy::cond::canceled`. Does nothing when not open. + */ + void close() noexcept; + + /** Check whether a descriptor is held. + + @return `true` if a descriptor is held and ready for I/O. + */ + bool is_open() const noexcept + { + return h_ && get().native_handle() >= 0; + } + + /** Get the native descriptor. + + @return The native descriptor, or -1 when not open. + */ + native_handle_type native_handle() const noexcept; + + /** Cancel pending asynchronous operations. + + Outstanding operations complete with a code that compares + equal to `capy::cond::canceled`. + */ + void cancel() noexcept; + + /** Wait for readiness without transferring bytes. + + Never reads, writes or modifies the descriptor -- including + its flags -- which is what makes it safe on a descriptor + another library owns. + + @param w The direction to wait on. + + @return An awaitable yielding `capy::io_result<>`. Yields + `errc::bad_file_descriptor` when not open. + + @par Example + @par !example wait + + @see wait_type + */ + [[nodiscard]] wait_awaitable wait(wait_type w) + { + return wait_awaitable(*this, w); + } + +protected: + /// Default-construct (for derived types that initialize `io_object` directly). + posix_descriptor() noexcept = default; + + /** Construct from a handle. + + @param h The handle this object takes ownership of. + */ + explicit posix_descriptor(handle h) noexcept : io_object(std::move(h)) {} + +private: + /// Return the implementation downcast to this type's interface. + implementation& get() const noexcept + { + return *static_cast(h_.get()); + } +}; + +} // namespace boost::corosio + +#endif // BOOST_COROSIO_POSIX || BOOST_COROSIO_MRDOCS + +#endif diff --git a/include/boost/corosio/random_access_file.hpp b/include/boost/corosio/random_access_file.hpp index 047600d8e..7f07f2cd0 100644 --- a/include/boost/corosio/random_access_file.hpp +++ b/include/boost/corosio/random_access_file.hpp @@ -399,19 +399,47 @@ class BOOST_COROSIO_DECL random_access_file : public io_object /** Adopt an existing native handle. - Closes any currently open file before adopting. - The file object takes ownership of the handle. Handles - created elsewhere may be unsuitable for asynchronous I/O; - such failures are reported through the returned error code. + Validation runs before anything is mutated or closed. On + error the object still holds whatever file it held before, + and the caller still owns @p handle. On success the object + takes ownership of @p handle and closes any file it + previously held. Handles created elsewhere may be unsuitable + for asynchronous I/O; such failures are reported through the + returned error code. @param handle The native file descriptor or handle. - @return The error code, empty on success. + @return An error code describing the outcome. The codes + that follow are those of the POSIX and io_uring + backends. `errc::invalid_argument` if @p handle is the + one this object already holds. + `errc::bad_file_descriptor` if it is invalid. + `errc::operation_not_supported` if it names something a + file object cannot position. Otherwise, the `errno` + reported by the kernel, or an empty code. The Windows + backend validates nothing and reports the Win32 error + from IOCP registration. + + @par Exception Safety + Throws nothing. Strong guarantee. + + @note On POSIX, "something a file object cannot position" + means, in practice, a pipe, a socket, or any other + anonymous inode. Adopt those into a @ref posix_descriptor + instead. + + @note The strong guarantee above holds on the POSIX and + io_uring backends. On Windows (IOCP), a failed @p handle + can still close the file this object held. That backend's + file services are unified in a later stage, which is + where this gap is closed. + + @see release */ [[nodiscard]] std::error_code assign(native_handle_type handle) noexcept; protected: - /// Construct from a pre-built handle (for native_random_access_file). + /// Construct from a pre-built handle (for `native_random_access_file`). explicit random_access_file(handle h) noexcept : io_object(std::move(h)) {} private: diff --git a/include/boost/corosio/stream_file.hpp b/include/boost/corosio/stream_file.hpp index 53cee16bc..bd559a3c3 100644 --- a/include/boost/corosio/stream_file.hpp +++ b/include/boost/corosio/stream_file.hpp @@ -277,14 +277,42 @@ class BOOST_COROSIO_DECL stream_file : public io_stream /** Adopt an existing native handle. - Closes any currently open file before adopting. - The file object takes ownership of the handle. Handles - created elsewhere may be unsuitable for asynchronous I/O. - Such failures are reported through the returned error code. + Validation runs before anything is mutated or closed. On + error the object still holds whatever file it held before, + and the caller still owns @p handle. On success the object + takes ownership of @p handle and closes any file it + previously held. Handles created elsewhere may be unsuitable + for asynchronous I/O; such failures are reported through the + returned error code. @param handle The native file descriptor or handle. - @return The error code, empty on success. + @return An error code describing the outcome. The codes + that follow are those of the POSIX and io_uring + backends. `errc::invalid_argument` if @p handle is the + one this object already holds. + `errc::bad_file_descriptor` if it is invalid. + `errc::operation_not_supported` if it names something a + file object cannot position. Otherwise, the `errno` + reported by the kernel, or an empty code. The Windows + backend validates nothing and reports the Win32 error + from IOCP registration. + + @par Exception Safety + Throws nothing. Strong guarantee. + + @note On POSIX, "something a file object cannot position" + means, in practice, a pipe, a socket, or any other + anonymous inode. Adopt those into a @ref posix_descriptor + instead. + + @note The strong guarantee above holds on the POSIX and + io_uring backends. On Windows (IOCP), a failed @p handle + can still close the file this object held. That backend's + file services are unified in a later stage, which is + where this gap is closed. + + @see release */ [[nodiscard]] std::error_code assign(native_handle_type handle) noexcept; @@ -305,10 +333,10 @@ class BOOST_COROSIO_DECL stream_file : public io_stream file_base::seek_basis origin = file_base::seek_set) noexcept; protected: - /// Default-construct (for derived types that initialize io_object directly). + /// Default-construct (for derived types that initialize `io_object` directly). stream_file() noexcept = default; - /** Construct from a pre-built handle (for native_stream_file). + /** Construct from a pre-built handle (for `native_stream_file`). @param h The pre-built handle to adopt. */ diff --git a/src/corosio/src/detail/use_backend_service.hpp b/src/corosio/src/detail/use_backend_service.hpp index 2f55f2188..a225884d7 100644 --- a/src/corosio/src/detail/use_backend_service.hpp +++ b/src/corosio/src/detail/use_backend_service.hpp @@ -42,6 +42,7 @@ #include #endif #if BOOST_COROSIO_HAS_URING +#include #include #include #include @@ -90,6 +91,12 @@ struct local_datagram_service_of using type = typename Tag::local_datagram_service_type; }; +template +struct descriptor_service_of +{ + using type = typename Tag::descriptor_service_type; +}; + template struct stream_file_service_of { diff --git a/src/corosio/src/posix_descriptor.cpp b/src/corosio/src/posix_descriptor.cpp new file mode 100644 index 000000000..a520d7ae9 --- /dev/null +++ b/src/corosio/src/posix_descriptor.cpp @@ -0,0 +1,80 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +#include + +#include + +#if BOOST_COROSIO_POSIX + +#include +#include + +#include "src/detail/use_backend_service.hpp" + +namespace boost::corosio { + +posix_descriptor::~posix_descriptor() +{ + close(); +} + +posix_descriptor::posix_descriptor(capy::execution_context& ctx) + : io_object(handle( + ctx, + detail::use_backend_service< + detail::descriptor_service_of, + detail::descriptor_service>(ctx))) +{ +} + +std::error_code +posix_descriptor::assign(native_handle_type fd) noexcept +{ + auto& svc = static_cast(h_.service()); + return svc.assign_descriptor(get(), fd); +} + +native_handle_type +posix_descriptor::release() +{ + if (!is_open()) + detail::throw_system_error( + make_error_code(std::errc::bad_file_descriptor), + "posix_descriptor::release"); + return get().release_descriptor(); +} + +void +posix_descriptor::close() noexcept +{ + if (!is_open()) + return; + h_.service().close(h_); +} + +native_handle_type +posix_descriptor::native_handle() const noexcept +{ + if (!h_) + return -1; + return get().native_handle(); +} + +void +posix_descriptor::cancel() noexcept +{ + if (!is_open()) + return; + get().cancel(); +} + +} // namespace boost::corosio + +#endif // BOOST_COROSIO_POSIX diff --git a/src/corosio/src/random_access_file.cpp b/src/corosio/src/random_access_file.cpp index b3eb5b35d..bc2c68cb9 100644 --- a/src/corosio/src/random_access_file.cpp +++ b/src/corosio/src/random_access_file.cpp @@ -127,8 +127,9 @@ random_access_file::release() std::error_code random_access_file::assign(native_handle_type handle) noexcept { - if (is_open()) - close(); + // Closing here first, unconditionally, would defeat the impl's own + // validate-before-mutate contract: get().assign() validates handle + // and only then closes whatever this object currently holds. return get().assign(handle); } diff --git a/src/corosio/src/stream_file.cpp b/src/corosio/src/stream_file.cpp index 1e2fbcb80..d03adaeed 100644 --- a/src/corosio/src/stream_file.cpp +++ b/src/corosio/src/stream_file.cpp @@ -127,8 +127,9 @@ stream_file::release() std::error_code stream_file::assign(native_handle_type handle) noexcept { - if (is_open()) - close(); + // Closing here first, unconditionally, would defeat the impl's own + // validate-before-mutate contract: get().assign() validates handle + // and only then closes whatever this object currently holds. return get().assign(handle); } diff --git a/test/doc/reference/native_posix_descriptor.record.cpp b/test/doc/reference/native_posix_descriptor.record.cpp new file mode 100644 index 000000000..a262fa1d9 --- /dev/null +++ b/test/doc/reference/native_posix_descriptor.record.cpp @@ -0,0 +1,69 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +// Reference example injected into +// include/boost/corosio/native/native_posix_descriptor.hpp's documentation +// for native_posix_descriptor, by doc/addons/extensions/reference-snippets.lua. +// The tagged region is what the reference renders; scaffolding stays outside +// the tags. +// +// native_posix_descriptor is a class template (`template`); +// the reference slug drops the template parameter, but the example must +// still name a concrete backend tag. corosio::epoll is what this library +// actually offers as a compile-time tag on Linux (see backend.hpp); every +// backend tag defines descriptor_type, so the type itself is not the +// constraint -- the tag's own existence is. Guarding on +// BOOST_COROSIO_HAS_EPOLL (rather than BOOST_COROSIO_POSIX) matches +// native_local_stream_socket.record.cpp's precedent for this exact class +// of example. + +#include "../doc_warnings.hpp" + +#include +#include +#include + +#include + +#if BOOST_COROSIO_HAS_EPOLL +#include +#endif + +namespace corosio = boost::corosio; +namespace capy = boost::capy; + +namespace { + +#if BOOST_COROSIO_HAS_EPOLL +// tag::assign_and_wait[] +capy::task +await_readable_native(int fd) +{ + corosio::native_io_context ctx; + corosio::native_posix_descriptor d(ctx); + + // Adopt a duplicate: assign() takes ownership, and the caller's fd + // must outlive it. + int copy = ::dup(fd); + if (copy < 0) + co_return std::error_code(errno, std::system_category()); + if (auto ec = d.assign(copy)) + { + // A rejected descriptor stays the caller's to close. + ::close(copy); + co_return ec; + } + + auto [ec] = co_await d.wait(corosio::wait_type::read); + co_return ec; +} +// end::assign_and_wait[] +#endif // BOOST_COROSIO_HAS_EPOLL + +} // namespace diff --git a/test/doc/reference/posix_descriptor__wait.function.cpp b/test/doc/reference/posix_descriptor__wait.function.cpp new file mode 100644 index 000000000..c62e8f077 --- /dev/null +++ b/test/doc/reference/posix_descriptor__wait.function.cpp @@ -0,0 +1,62 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +// Reference example injected into include/boost/corosio/posix_descriptor.hpp's +// documentation for posix_descriptor::wait, by +// doc/addons/extensions/reference-snippets.lua. The tagged region is what the +// reference renders; scaffolding stays outside the tags. +// +// posix_descriptor's whole class body is wrapped in #if BOOST_COROSIO_POSIX in +// its own header, so the region that names the type is guarded the same way, +// following local_datagram_socket.record.cpp. The includes below are safe +// unconditionally -- the header itself resolves to nothing off POSIX. + +#include "../doc_warnings.hpp" + +#include +#include +#include + +#include + +#if BOOST_COROSIO_POSIX +#include +#endif + +namespace corosio = boost::corosio; +namespace capy = boost::capy; + +namespace { + +#if BOOST_COROSIO_POSIX +// tag::wait[] +capy::task +await_readable(corosio::io_context& ioc, int fd) +{ + // Adopt a duplicate: assign() takes ownership, and the caller's fd + // must outlive it. + int copy = ::dup(fd); + if (copy < 0) + co_return std::error_code(errno, std::system_category()); + + corosio::posix_descriptor d(ioc); + if (auto ec = d.assign(copy)) + { + // A rejected descriptor stays the caller's to close. + ::close(copy); + co_return ec; + } + + auto [ec] = co_await d.wait(corosio::wait_type::read); + co_return ec; +} +// end::wait[] +#endif + +} // namespace diff --git a/test/doc/snippets/4s_native_descriptors.cpp b/test/doc/snippets/4s_native_descriptors.cpp new file mode 100644 index 000000000..4afb8c021 --- /dev/null +++ b/test/doc/snippets/4s_native_descriptors.cpp @@ -0,0 +1,273 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +// Compiled fragments shown in pages/4.guide/4s.native-descriptors.adoc. + +// Fragments deliberately leave results and bindings unused; the pages +// explain the values in prose instead. +#if defined(__GNUC__) || defined(__clang__) +#pragma GCC diagnostic ignored "-Wunused-but-set-variable" +#pragma GCC diagnostic ignored "-Wunused-variable" +#pragma GCC diagnostic ignored "-Wunused-parameter" +#pragma GCC diagnostic ignored "-Wunused-value" +#pragma GCC diagnostic ignored "-Wunused-result" +#pragma GCC diagnostic ignored "-Wunused-function" +#endif +#if defined(_MSC_VER) +#pragma warning(disable : 4834) // discarding [[nodiscard]] return value +#pragma warning(disable : 4189) // local variable initialized but not referenced +#pragma warning(disable : 4100) // unreferenced formal parameter +#pragma warning(disable : 4101) // unreferenced local variable +#endif + +#include + +#include "test_suite.hpp" + +// The page documents a POSIX-only type, so the whole body is guarded -- +// including the `assume` fragment, which names . +#if BOOST_COROSIO_POSIX + +// tag::assume[] +#include +#include +#include +#include +#include +#include + +#include +#include + +#include + +namespace corosio = boost::corosio; +namespace capy = boost::capy; +// end::assume[] + +#include + +#include +#include + +#if defined(__linux__) +#include +#include +#endif + +namespace { + +std::error_code +last_error() noexcept +{ + return std::error_code(errno, std::system_category()); +} + +// tag::layering[] +// Nothing below is descriptor-specific. `capy::read` is constrained on +// capy::Stream, and a posix_descriptor models it exactly as a +// tcp_socket or a tls_stream does, so the same algorithm drives all +// three. +capy::task +fill(corosio::posix_descriptor& d, capy::mutable_buffer buf) +{ + auto [ec, n] = co_await capy::read(d, buf); + co_return ec; +} +// end::layering[] + +#if defined(__linux__) + +capy::task +count_events(corosio::io_context& ioc) +{ + // tag::adopt_eventfd[] + int fd = ::eventfd(0, EFD_CLOEXEC); + if (fd < 0) + co_return last_error(); + + corosio::posix_descriptor d(ioc); + if (auto ec = d.assign(fd)) + { + // A failed assign() leaves the descriptor with the caller. + ::close(fd); + co_return ec; + } + + // An eventfd delivers its accumulated count as a single 8-byte + // host-order integer; the read parks until the count is nonzero + // and resets it to zero. + std::uint64_t count = 0; + auto [ec, n] = + co_await d.read_some(capy::mutable_buffer(&count, sizeof(count))); + // end::adopt_eventfd[] + co_return ec; +} + +capy::task +watch_directory(corosio::io_context& ioc, char const* path) +{ + // tag::adopt_inotify[] + int fd = ::inotify_init1(IN_CLOEXEC); + if (fd < 0) + co_return last_error(); + if (::inotify_add_watch(fd, path, IN_CREATE | IN_DELETE) < 0) + { + auto ec = last_error(); + ::close(fd); + co_return ec; + } + + corosio::posix_descriptor d(ioc); + if (auto ec = d.assign(fd)) + { + ::close(fd); + co_return ec; + } + + // inotify delivers whole events. The buffer must be aligned for + // inotify_event and large enough for at least one event plus its + // variable-length name. + alignas(struct inotify_event) char buf[4096]; + auto [ec, n] = co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + if (ec) + co_return ec; + + auto const* ev = reinterpret_cast(buf); + // end::adopt_inotify[] + (void)ev; + co_return ec; +} + +#endif // __linux__ + +capy::task +read_a_line(corosio::io_context& ioc) +{ + // tag::wait_only[] + // wait() transfers no bytes and sets no flag, so standard input -- + // and the terminal the parent shell shares with it -- stays exactly + // as the process inherited it. + int fd = ::dup(STDIN_FILENO); + if (fd < 0) + co_return last_error(); + + corosio::posix_descriptor d(ioc); + if (auto ec = d.assign(fd)) + { + ::close(fd); + co_return ec; + } + + auto [ec] = co_await d.wait(corosio::wait_type::read); + if (ec) + co_return ec; + + // Readable, so a blocking ::read returns here. Readiness is not a + // general guarantee against parking -- a socket can report ready + // and still have nothing to hand over -- but it holds for a tty. + char line[256]; + auto n = ::read(d.native_handle(), line, sizeof(line)); + // end::wait_only[] + if (n < 0) + co_return last_error(); + co_return std::error_code{}; +} + +// Stand-in for a C library that owns a descriptor and does its own I/O +// on it (the libpq shape). +struct foreign_conn +{}; + +int +foreign_fd(foreign_conn*) noexcept +{ + return -1; +} + +capy::task +watch_foreign(corosio::io_context& ioc, foreign_conn* conn) +{ + // tag::dup_for_foreign_fd[] + // Adopt a duplicate, never the library's own descriptor. Both refer + // to one open file description, so readiness is identical and the + // lazily applied O_NONBLOCK is visible to the library either way -- + // but corosio's close() can only ever close the copy. + int copy = ::dup(foreign_fd(conn)); + if (copy < 0) + co_return last_error(); + + corosio::posix_descriptor d(ioc); + if (auto ec = d.assign(copy)) + { + ::close(copy); + co_return ec; + } + // end::dup_for_foreign_fd[] + + auto [ec] = co_await d.wait(corosio::wait_type::read); + co_return ec; +} + +capy::task<> +run_and_store(capy::task t, std::error_code& ec_out) +{ + ec_out = co_await std::move(t); +} + +struct native_descriptors_test +{ + // Adopt the read end of a pipe and drive it through the generic + // stream algorithm: the claim the page makes about layering is the + // one worth testing for real. + void testLayering() + { + int fds[2]; + BOOST_TEST(::pipe(fds) == 0); + + corosio::io_context ioc; + corosio::posix_descriptor d(ioc); + BOOST_TEST(!d.assign(fds[0])); + + char got[5] = {}; + std::error_code ec = std::make_error_code(std::errc::io_error); + capy::run_async(ioc.get_executor())( + run_and_store(fill(d, capy::mutable_buffer(got, sizeof(got))), ec)); + + // All five bytes are in the pipe before the read issues, so + // this checks the adopt -> capy::read -> read_some path end to + // end, not the algorithm's short-read branch. + BOOST_TEST(::write(fds[1], "hello", 5) == 5); + + ioc.run(); + BOOST_TEST(!ec); + BOOST_TEST(std::memcmp(got, "hello", 5) == 0); + ::close(fds[1]); + } + + void run() + { + testLayering(); + } +}; + +} // namespace + +#else + +namespace { +struct native_descriptors_test +{ + void run() {} +}; +} // namespace + +#endif // BOOST_COROSIO_POSIX + +TEST_SUITE(native_descriptors_test, "boost.corosio.doc.4s_native_descriptors"); diff --git a/test/unit/lazy_services.cpp b/test/unit/lazy_services.cpp index 0ccc6d802..12e65e6be 100644 --- a/test/unit/lazy_services.cpp +++ b/test/unit/lazy_services.cpp @@ -17,6 +17,7 @@ #include #include #include +#include #include #include #include @@ -38,6 +39,7 @@ #include #include #else +#include #include #include #include @@ -82,6 +84,7 @@ using local_stream_service_t = detail::local_stream_service; using local_stream_acceptor_service_t = detail::local_stream_acceptor_service; using local_datagram_service_t = detail::local_datagram_service; +using descriptor_service_t = detail::descriptor_service; using resolver_service_t = detail::posix_resolver_service; using signal_service_t = detail::posix_signal_service; using file_service_t = detail::file_service; @@ -109,6 +112,8 @@ struct lazy_services_test #if !BOOST_COROSIO_HAS_IOCP BOOST_TEST( ioc.template find_service() == nullptr); + BOOST_TEST( + ioc.template find_service() == nullptr); #endif BOOST_TEST(ioc.template find_service() == nullptr); BOOST_TEST(ioc.template find_service() == nullptr); @@ -157,6 +162,10 @@ struct lazy_services_test local_datagram_socket ld(ioc); BOOST_TEST( ioc.template find_service() != nullptr); + + posix_descriptor pd(ioc); + BOOST_TEST( + ioc.template find_service() != nullptr); #endif resolver res(ioc); diff --git a/test/unit/native/native_posix_descriptor.cpp b/test/unit/native/native_posix_descriptor.cpp new file mode 100644 index 000000000..36144b3d5 --- /dev/null +++ b/test/unit/native/native_posix_descriptor.cpp @@ -0,0 +1,156 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +#include + +#include + +#if BOOST_COROSIO_POSIX + +#include + +#include +#include +#include + +#include +#include +#include + +#include + +#include "context.hpp" +#include "test_suite.hpp" + +namespace boost::corosio { + +template +struct native_posix_descriptor_test +{ + static_assert( + std::is_base_of_v>); + + static_assert( + !std::is_same_v< + decltype(std::declval&>() + .read_some(std::declval())), + decltype(std::declval().read_some( + std::declval()))>, + "native_posix_descriptor::read_some must shadow " + "io_stream::read_some"); + static_assert( + !std::is_same_v< + decltype(std::declval&>() + .write_some(std::declval())), + decltype(std::declval().write_some( + std::declval()))>, + "native_posix_descriptor::write_some must shadow " + "io_stream::write_some"); + static_assert( + !std::is_same_v< + decltype(std::declval&>().wait( + wait_type::read)), + decltype(std::declval().wait(wait_type::read))>, + "native_posix_descriptor::wait must shadow posix_descriptor::wait"); + + void testConstruct() + { + native_io_context ioc; + native_posix_descriptor d(ioc); + BOOST_TEST_EQ(d.is_open(), false); + } + + // Writes into a pipe with plain POSIX calls and reads the bytes + // back through native_posix_descriptor, exercising the shadowed + // read_some() awaitable end to end rather than merely + // static-asserting the type shape. + void testReadWriteRoundTrip() + { + native_io_context ioc; + native_posix_descriptor d(ioc); + + int fds[2]; + BOOST_TEST_EQ(::pipe(fds), 0); + BOOST_TEST(!d.assign(fds[0])); + + char const msg[] = "abc"; + BOOST_TEST_EQ(::write(fds[1], msg, 3), 3); + + char buf[8] = {}; + std::size_t n = 0; + std::error_code ec; + auto reader = [&]() -> capy::task<> { + auto [rec, rn] = + co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + ec = rec; + n = rn; + }; + capy::run_async(ioc.get_executor())(reader()); + ioc.run(); + + BOOST_TEST_EQ(ec, std::error_code{}); + BOOST_TEST_EQ(n, 3u); + BOOST_TEST_EQ(std::string(buf, n), std::string("abc")); + + ::close(fds[1]); + } + + void testPolymorphicSlice() + { + native_io_context ioc; + native_posix_descriptor d(ioc); + posix_descriptor& base = d; + BOOST_TEST_EQ(base.is_open(), false); + } + + // Exercises the shadowed wait() awaitable through a genuine + // round trip: the reader parks until the writer produces data. + void testWait() + { + native_io_context ioc; + native_posix_descriptor d(ioc); + + int fds[2]; + BOOST_TEST_EQ(::pipe(fds), 0); + BOOST_TEST(!d.assign(fds[0])); + + char const msg[] = "x"; + BOOST_TEST_EQ(::write(fds[1], msg, 1), 1); + + std::error_code ec; + bool done = false; + auto waiter = [&]() -> capy::task<> { + auto [wec] = co_await d.wait(wait_type::read); + ec = wec; + done = true; + }; + capy::run_async(ioc.get_executor())(waiter()); + ioc.run(); + + BOOST_TEST(done); + BOOST_TEST_EQ(ec, std::error_code{}); + + ::close(fds[1]); + } + + void run() + { + testConstruct(); + testReadWriteRoundTrip(); + testPolymorphicSlice(); + testWait(); + } +}; + +COROSIO_BACKEND_TESTS( + native_posix_descriptor_test, "boost.corosio.native_posix_descriptor") + +} // namespace boost::corosio + +#endif // BOOST_COROSIO_POSIX diff --git a/test/unit/native/native_resume_cancel.cpp b/test/unit/native/native_resume_cancel.cpp index 00d962ffa..9e4986827 100644 --- a/test/unit/native/native_resume_cancel.cpp +++ b/test/unit/native/native_resume_cancel.cpp @@ -17,6 +17,7 @@ #include #include +#include #include #include #include @@ -42,6 +43,8 @@ #include #include #include + +#include #endif #include "context.hpp" @@ -401,6 +404,96 @@ struct native_resume_cancel_test BOOST_TEST_EQ(canceled, 4); } + void testDescriptorPreStoppedPerformsNoIo() + { + // The pre-stopped half of the contract, with the stronger + // assertion: not merely that the operation reports canceled, + // but that no read reached the descriptor. An awaitable that + // checks the token only at resume lets a speculative ::read() + // drain the pipe first and then discards the bytes, which is + // silent data loss rather than a cancellation. + native_io_context ioc; + auto ex = ioc.get_executor(); + + int fds[2]; + BOOST_TEST_EQ(::pipe(fds), 0); + BOOST_TEST_EQ(::write(fds[1], "hello", 5), 5); + + native_posix_descriptor d(ioc); + BOOST_TEST(!d.assign(fds[0])); + + std::stop_source ss; + ss.request_stop(); + + char buf[8] = {}; + int canceled = 0; + auto driver = [&]() -> capy::task<> { + auto [rec, rn] = + co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + if (rec == capy::cond::canceled && rn == 0) + ++canceled; + auto [tec] = co_await d.wait(wait_type::read); + if (tec == capy::cond::canceled) + ++canceled; + }; + capy::run_async(ex, ss.get_token())(driver()); + ioc.run(); + BOOST_TEST_EQ(canceled, 2); + + // The payload must still be in the pipe: a cancelled read + // consumes nothing. + char check[8] = {}; + BOOST_TEST_EQ(::read(d.native_handle(), check, sizeof(check)), 5); + BOOST_TEST(std::memcmp(check, "hello", 5) == 0); + + ::close(fds[1]); + } + + void testDescriptorStopAfterCompleted() + { + // The racing half: a stop that lands on an already-completed + // transfer changes nothing, and the byte count survives. + io_context_options opts; + opts.inline_budget_max = 0; + native_io_context ioc(opts); + auto ex = ioc.get_executor(); + + int fds[2]; + BOOST_TEST_EQ(::pipe(fds), 0); + BOOST_TEST_EQ(::write(fds[1], "hello", 5), 5); + + native_posix_descriptor d(ioc); + BOOST_TEST(!d.assign(fds[0])); + + std::stop_source ss; + std::error_code rec = capy::error::eof; + std::size_t rn = 0; + char buf[8] = {}; + bool done = false; + + auto reader = [&]() -> capy::task<> { + auto [ec, n] = + co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + rec = ec; + rn = n; + done = true; + }; + auto stopper = [&]() -> capy::task<> { + ss.request_stop(); + co_return; + }; + capy::run_async(ex, ss.get_token())(reader()); + capy::run_async(ex)(stopper()); + ioc.run(); + + BOOST_TEST(done); + BOOST_TEST(!rec); + BOOST_TEST_EQ(rn, 5u); + BOOST_TEST(std::memcmp(buf, "hello", 5) == 0); + + ::close(fds[1]); + } + void testLocalStreamAcceptorPreStopped() { native_io_context ioc; @@ -503,6 +596,8 @@ struct native_resume_cancel_test testFileResumeCancel(); #if BOOST_COROSIO_POSIX testLocalStreamPreStopped(); + testDescriptorPreStoppedPerformsNoIo(); + testDescriptorStopAfterCompleted(); testLocalStreamAcceptorPreStopped(); testLocalDatagramPreStopped(); #endif diff --git a/test/unit/native/uring/descriptor_continue.cpp b/test/unit/native/uring/descriptor_continue.cpp new file mode 100644 index 000000000..2a5a52ae5 --- /dev/null +++ b/test/unit/native/uring/descriptor_continue.cpp @@ -0,0 +1,302 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +// The two-phase transfer machinery in uring_descriptor: a kernel EAGAIN +// arms a poll_add, the poll's completion re-submits the transfer, and +// every path that abandons the op in between leaves a terminal result +// behind. None of this is reachable through posix_descriptor's public +// API: io_uring retries a pollable O_NONBLOCK descriptor internally, so +// -EAGAIN never reaches userspace for the descriptor kinds this type +// carries, and the remaining producer (an SQ ring that stays full after +// a flush) needs a CQ overflow that a unit test cannot force on a +// 256-entry ring. The CQE is therefore injected here and the rest of +// the path -- the real op types, the real ring, a real pipe -- runs +// unmodified. + +#include "test_suite.hpp" + +#include + +#if BOOST_COROSIO_HAS_URING + +#include +#include + +#include +#include + +#include +#include +#include +#include + +#include +#include +#include + +namespace boost::corosio { + +// Exposes the io_uring scheduler, the way multishot_acceptor.cpp does. +struct uring_descriptor_test_context : native_io_context +{ + detail::uring_scheduler& scheduler() noexcept + { + return *static_cast(sched_); + } +}; + +// A real coroutine handle for the completing op to resume. The body runs +// on the first resume, because initial_suspend suspends. +struct probe_coro +{ + struct promise_type + { + probe_coro get_return_object() noexcept + { + return {std::coroutine_handle::from_promise(*this)}; + } + std::suspend_always initial_suspend() noexcept + { + return {}; + } + std::suspend_always final_suspend() noexcept + { + return {}; + } + void return_void() noexcept {} + void unhandled_exception() noexcept {} + }; + + std::coroutine_handle h; +}; + +inline probe_coro +make_probe_coro(bool& resumed) +{ + resumed = true; + co_return; +} + +struct uring_descriptor_continue_test +{ + // What uring_descriptor::arm_slot does at submission time; the + // epochs it snapshots are the whole subject of these tests. + template + static void arm(detail::uring_descriptor& d, Op& op) + { + op.desc = &d; + op.polling = false; + op.cancel_epoch = d.cancel_epoch(); + op.desc_epoch = d.desc_epoch(); + } + + // The handler's tail, so an assertion reads the same two values a + // caller of read_some() would. + static void + decode(detail::uring_op& op, std::error_code& ec, std::size_t& bytes) + { + op.ec_out = &ec; + detail::uring_set_result(&op, /*is_read=*/true, op.empty_buffer); + bytes = op.res >= 0 ? static_cast(op.res) : 0u; + } + + void testCancelledMidPollDoesNotReportRevents() + { + // A stop_token firing between a successful poll CQE and its + // dispatch. `res` still holds the revents mask, which the + // completion decode would read as a byte count. + uring_descriptor_test_context ctx; + auto d = std::make_shared(ctx.scheduler()); + + detail::uring_descriptor_read_op op; + arm(*d, op); + op.polling = true; + op.res = POLLIN; + op.cancelled.store(true, std::memory_order_release); + + BOOST_TEST(!detail::uring_descriptor_continue(op)); + BOOST_TEST_EQ(op.res, -ECANCELED); + + std::error_code ec; + std::size_t bytes = 99; + decode(op, ec, bytes); + BOOST_TEST(ec == capy::cond::canceled); + // Without the terminal result this is POLLIN, i.e. one byte the + // caller never read. + BOOST_TEST_EQ(bytes, 0u); + } + + void testCancelInTheGapStopsTheRearm() + { + // cancel() lands while the op is out of the ring, so its + // cancel-by-fd SQE finds nothing and the op's own `cancelled` + // flag is never set. The epoch is the only thing that sees it. + uring_descriptor_test_context ctx; + auto d = std::make_shared(ctx.scheduler()); + + detail::uring_descriptor_write_op op; + arm(*d, op); + op.polling = true; + op.res = POLLOUT; + + BOOST_TEST(!op.cancelled.load(std::memory_order_acquire)); + d->cancel(); + BOOST_TEST(!op.cancelled.load(std::memory_order_acquire)); + + BOOST_TEST(!detail::uring_descriptor_continue(op)); + BOOST_TEST_EQ(op.res, -ECANCELED); + } + + void testEagainCancelInTheGapDoesNotPark() + { + // The hang the epoch prevents: a transfer that answered EAGAIN + // would otherwise arm a fresh poll nothing can cancel, and + // run() would never return. + uring_descriptor_test_context ctx; + auto d = std::make_shared(ctx.scheduler()); + + detail::uring_descriptor_read_op op; + arm(*d, op); + op.res = -EAGAIN; + d->cancel(); + + BOOST_TEST(!detail::uring_descriptor_continue(op)); + // Still the transfer phase: nothing was submitted, so nothing + // is parked. + BOOST_TEST(!op.polling); + BOOST_TEST_EQ(op.res, -EAGAIN); + } + + void testRecycledFdNumberStopsTheRearm() + { + // The aliasing case a bare native_handle() == op.fd comparison + // cannot see: the descriptor is re-adopted and lands on the + // same fd number. + int fds[2]; + BOOST_TEST_EQ(::pipe(fds), 0); + + uring_descriptor_test_context ctx; + auto d = std::make_shared(ctx.scheduler()); + d->set_descriptor(fds[0]); + + detail::uring_descriptor_read_op op; + arm(*d, op); + op.fd = fds[0]; + op.polling = true; + op.res = POLLIN; + + d->set_descriptor(fds[0]); + BOOST_TEST_EQ(d->native_handle(), op.fd); // the fd check would pass + + BOOST_TEST(!detail::uring_descriptor_continue(op)); + BOOST_TEST_EQ(op.res, -EBADF); + + std::error_code ec; + std::size_t bytes = 99; + decode(op, ec, bytes); + // Without the terminal result: no error at all, and a + // one-byte read of an untouched buffer. + BOOST_TEST(ec == std::errc::bad_file_descriptor); + BOOST_TEST_EQ(bytes, 0u); + + ::close(fds[1]); + } + + void testTerminalResultsAreLeftAlone() + { + uring_descriptor_test_context ctx; + auto d = std::make_shared(ctx.scheduler()); + + // A poll that failed or was cancelled is already the answer. + detail::uring_descriptor_read_op poll_failed; + arm(*d, poll_failed); + poll_failed.polling = true; + poll_failed.res = -ECANCELED; + BOOST_TEST(!detail::uring_descriptor_continue(poll_failed)); + BOOST_TEST_EQ(poll_failed.res, -ECANCELED); + + // A transfer error is not a readiness mask; leave it. + detail::uring_descriptor_write_op failed; + arm(*d, failed); + failed.res = -EPIPE; + BOOST_TEST(!detail::uring_descriptor_continue(failed)); + BOOST_TEST_EQ(failed.res, -EPIPE); + + // Nor is a genuine byte count. + detail::uring_descriptor_read_op transferred; + arm(*d, transferred); + transferred.res = 5; + BOOST_TEST(!detail::uring_descriptor_continue(transferred)); + BOOST_TEST_EQ(transferred.res, 5); + } + + void testEagainArmsPollThenRetriesTransfer() + { + // The forward path, end to end on a real ring: injected EAGAIN + // -> poll_add -> readiness -> re-submitted READV -> coroutine + // resumed with the bytes. + int fds[2]; + BOOST_TEST_EQ(::pipe(fds), 0); + + uring_descriptor_test_context ctx; + auto ex = ctx.get_executor(); + auto d = std::make_shared(ctx.scheduler()); + d->set_descriptor(fds[0]); + + bool resumed = false; + auto coro = make_probe_coro(resumed); + + char buf[8]{}; + capy::mutable_buffer mb(buf, sizeof(buf)); + std::error_code ec; + std::size_t bytes = 0; + + detail::uring_descriptor_read_op op; + op.prepare( + coro.h, ex, &ec, &bytes, fds[0], /*file_offset=*/-1, + &ctx.scheduler(), d, mb, std::stop_token{}); + arm(*d, op); + + // Data is already waiting, so the poll this arms fires at once. + BOOST_TEST_EQ(::write(fds[1], "hi", 2), 2); + + op.res = -EAGAIN; + BOOST_TEST(detail::uring_descriptor_continue(op)); + BOOST_TEST(op.polling); // phase flipped; a poll_add is in the ring + + ctx.run(); + + BOOST_TEST(resumed); + BOOST_TEST(!op.polling); // the poll's CQE re-submitted the read + BOOST_TEST(!ec); + BOOST_TEST_EQ(bytes, 2u); + BOOST_TEST_EQ(std::memcmp(buf, "hi", 2), 0); + + coro.h.destroy(); + ::close(fds[1]); + } + + void run() + { + testCancelledMidPollDoesNotReportRevents(); + testCancelInTheGapStopsTheRearm(); + testEagainCancelInTheGapDoesNotPark(); + testRecycledFdNumberStopsTheRearm(); + testTerminalResultsAreLeftAlone(); + testEagainArmsPollThenRetriesTransfer(); + } +}; + +TEST_SUITE( + uring_descriptor_continue_test, + "boost.corosio.native.uring.descriptor_continue"); + +} // namespace boost::corosio + +#endif // BOOST_COROSIO_HAS_URING diff --git a/test/unit/native/validate_fd.cpp b/test/unit/native/validate_fd.cpp new file mode 100644 index 000000000..0437dbad8 --- /dev/null +++ b/test/unit/native/validate_fd.cpp @@ -0,0 +1,169 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +// Test that header file is self-contained. +#include + +#include + +#if BOOST_COROSIO_POSIX + +#include +#include + +#include +#include +#include + +#if defined(__linux__) +#include +#endif + +#include "test_suite.hpp" + +namespace boost::corosio { + +struct validate_fd_test +{ + void testRejectsBadFd() + { + BOOST_TEST( + detail::validate_descriptor_fd(-1) == + std::errc::bad_file_descriptor); + + int fds[2]; + BOOST_TEST_EQ(::pipe(fds), 0); + ::close(fds[0]); + ::close(fds[1]); + BOOST_TEST( + detail::validate_descriptor_fd(fds[0]) == + std::errc::bad_file_descriptor); + } + + void testAcceptsPipe() + { + int fds[2]; + BOOST_TEST_EQ(::pipe(fds), 0); + BOOST_TEST(!detail::validate_descriptor_fd(fds[0])); + BOOST_TEST(!detail::validate_descriptor_fd(fds[1])); + ::close(fds[0]); + ::close(fds[1]); + } + + void testAcceptsAnonymousInode() + { + // The flagship fd kinds (eventfd, timerfd, inotify, pidfd) are + // anonymous inodes whose st_mode type bits are all zero. This is + // exactly why the check is a reject-list and not an accept-list. + // + // /dev/null stands in for a character device here because it is + // portable, but note that passing this check is not a promise + // that assign() will succeed: /dev/null is not pollable, so + // epoll refuses it with EPERM and kqueue with EINVAL. The + // validator's job is the file-type policy, not reachability. + int fd = ::open("/dev/null", O_RDWR | O_CLOEXEC); + BOOST_TEST(fd >= 0); + BOOST_TEST(!detail::validate_descriptor_fd(fd)); + ::close(fd); + } + + void testRejectsRegularFile() + { + // stream_file::assign() / random_access_file::assign() own + // adoption of regular files; posix_descriptor rejects by policy. + auto path = + std::filesystem::temp_directory_path() / "corosio_validate_fd_test"; + int fd = ::open(path.c_str(), O_RDWR | O_CREAT | O_CLOEXEC, 0600); + BOOST_TEST(fd >= 0); + BOOST_TEST( + detail::validate_descriptor_fd(fd) == + std::errc::operation_not_supported); + ::close(fd); + std::filesystem::remove(path); + } + + void testRejectsDirectory() + { + int fd = ::open(".", O_RDONLY | O_DIRECTORY | O_CLOEXEC); + BOOST_TEST(fd >= 0); + BOOST_TEST( + detail::validate_descriptor_fd(fd) == + std::errc::operation_not_supported); + ::close(fd); + } + + void testFileFdAcceptListInversion() + { + // validate_file_fd is an accept-list, the deliberate inverse of + // validate_descriptor_fd's reject-list: a regular file is the + // fd kind a file object exists to adopt. + auto path = std::filesystem::temp_directory_path() / + "corosio_validate_file_fd_test"; + int fd = ::open(path.c_str(), O_RDWR | O_CREAT | O_CLOEXEC, 0600); + BOOST_TEST(fd >= 0); + BOOST_TEST(!detail::validate_file_fd(fd)); + ::close(fd); + std::filesystem::remove(path); + +#if defined(__linux__) + // Head-to-head on the same fd: an anonymous inode (no st_mode + // type bits) is exactly what validate_descriptor_fd exists to + // admit and exactly what validate_file_fd -- needing a + // positionable fd -- must reject. + int efd = ::eventfd(0, EFD_CLOEXEC); + BOOST_TEST(efd >= 0); + BOOST_TEST(!detail::validate_descriptor_fd(efd)); + BOOST_TEST( + detail::validate_file_fd(efd) == + std::errc::operation_not_supported); + ::close(efd); +#endif + } + + void testEnsureNonblockingIsIdempotentAndLazy() + { + int fds[2]; + BOOST_TEST_EQ(::pipe(fds), 0); + + // Validation alone must never touch the flags. + BOOST_TEST(!detail::validate_descriptor_fd(fds[0])); + BOOST_TEST_EQ(::fcntl(fds[0], F_GETFL) & O_NONBLOCK, 0); + + BOOST_TEST(!detail::ensure_nonblocking(fds[0])); + BOOST_TEST(::fcntl(fds[0], F_GETFL) & O_NONBLOCK); + + // Second call is a no-op and still succeeds. + BOOST_TEST(!detail::ensure_nonblocking(fds[0])); + BOOST_TEST(::fcntl(fds[0], F_GETFL) & O_NONBLOCK); + + ::close(fds[0]); + ::close(fds[1]); + + BOOST_TEST( + detail::ensure_nonblocking(fds[0]) == + std::errc::bad_file_descriptor); + } + + void run() + { + testRejectsBadFd(); + testAcceptsPipe(); + testAcceptsAnonymousInode(); + testRejectsRegularFile(); + testRejectsDirectory(); + testFileFdAcceptListInversion(); + testEnsureNonblockingIsIdempotentAndLazy(); + } +}; + +TEST_SUITE(validate_fd_test, "boost.corosio.validate_fd"); + +} // namespace boost::corosio + +#endif // BOOST_COROSIO_POSIX diff --git a/test/unit/posix_descriptor.cpp b/test/unit/posix_descriptor.cpp new file mode 100644 index 000000000..ffe7c67d7 --- /dev/null +++ b/test/unit/posix_descriptor.cpp @@ -0,0 +1,796 @@ +// +// Copyright (c) 2026 Michael Vandeberg +// +// Distributed under the Boost Software License, Version 1.0. (See accompanying +// file LICENSE_1_0.txt or copy at http://www.boost.org/LICENSE_1_0.txt) +// +// Official repository: https://github.com/cppalliance/corosio +// + +// Test that header file is self-contained. +#include + +#include + +#if BOOST_COROSIO_POSIX + +#include +#include +#include +#include +#include +#include +#include + +#include +#include +#include +#include +#include +#include +#include + +#include +#include + +#include "context.hpp" +#include "test_suite.hpp" + +namespace boost::corosio { + +static_assert(capy::ReadStream); +static_assert(capy::WriteStream); +static_assert(std::is_base_of_v); +static_assert(!std::is_copy_constructible_v); +static_assert(std::is_move_constructible_v); + +// assign() must be impossible to call and ignore. +static_assert(std::is_same_v< + decltype(std::declval().assign( + std::declval())), + std::error_code>); + +struct posix_descriptor_contract_test +{ + void run() {} +}; + +TEST_SUITE( + posix_descriptor_contract_test, "boost.corosio.posix_descriptor_contract"); + +template +struct posix_descriptor_test +{ + // Returns a pipe whose ends are both blocking, so every test that + // cares about O_NONBLOCK starts from a known state. + static void make_pipe(int (&fds)[2]) + { + BOOST_TEST_EQ(::pipe(fds), 0); + } + + // The .epoll, .select and .uring variants are separate ctest + // entries that run concurrently under ctest -j, so a fixed global + // name would race between them. + static std::filesystem::path temp_path(char const* tag) + { + return std::filesystem::temp_directory_path() / + ("corosio_posix_descriptor_" + std::string(tag) + "_" + + typeid(decltype(Backend)).name() + "_" + + std::to_string(::getpid())); + } + + void testConstruction() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + BOOST_TEST_EQ(d.is_open(), false); + BOOST_TEST_EQ(d.native_handle(), -1); + } + + void testAssignRejectsRegularFile() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + + auto path = temp_path("regular"); + int fd = ::open(path.c_str(), O_RDWR | O_CREAT | O_CLOEXEC, 0600); + BOOST_TEST(fd >= 0); + + BOOST_TEST(d.assign(fd) == std::errc::operation_not_supported); + // Rejection leaves the object closed and the fd with the caller. + BOOST_TEST_EQ(d.is_open(), false); + BOOST_TEST_EQ(::fcntl(fd, F_GETFD) >= 0, true); + + ::close(fd); + std::filesystem::remove(path); + } + + void testAssignRejectsBadFd() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + BOOST_TEST(d.assign(-1) == std::errc::bad_file_descriptor); + } + + void testAssignRejectsSelf() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + + BOOST_TEST(!d.assign(fds[0])); + BOOST_TEST(d.assign(d.native_handle()) == std::errc::invalid_argument); + // Still open on the same fd after the rejected self-assign. + BOOST_TEST_EQ(d.is_open(), true); + BOOST_TEST_EQ(d.native_handle(), fds[0]); + + ::close(fds[1]); + } + + void testFailedAssignLeavesPriorStateIntact() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[0])); + + auto path = temp_path("reject"); + int reg = ::open(path.c_str(), O_RDWR | O_CREAT | O_CLOEXEC, 0600); + BOOST_TEST(reg >= 0); + + BOOST_TEST(d.assign(reg) == std::errc::operation_not_supported); + + // The held descriptor survived the rejected assign and still works. + BOOST_TEST_EQ(d.native_handle(), fds[0]); + char const msg[] = "ok"; + BOOST_TEST_EQ(::write(fds[1], msg, 2), 2); + + auto ex = ioc.get_executor(); + char buf[8]{}; + std::size_t n = 0; + std::error_code ec; + auto reader = [&]() -> capy::task<> { + auto [rec, rn] = + co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + ec = rec; + n = rn; + }; + capy::run_async(ex)(reader()); + ioc.run(); + + BOOST_TEST(!ec); + BOOST_TEST_EQ(n, 2u); + + ::close(reg); + std::filesystem::remove(path); + ::close(fds[1]); + } + + void testAssignDoesNotTouchFlagsButReadDoes() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + + BOOST_TEST(!d.assign(fds[0])); + // The single most important contract: adoption is non-mutating. + BOOST_TEST_EQ(::fcntl(fds[0], F_GETFL) & O_NONBLOCK, 0); + + char const msg[] = "hello"; + BOOST_TEST_EQ(::write(fds[1], msg, 5), 5); + + char buf[16]{}; + std::size_t n = 0; + auto reader = [&]() -> capy::task<> { + auto [rec, rn] = + co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + BOOST_TEST(!rec); + n = rn; + }; + capy::run_async(ioc.get_executor())(reader()); + ioc.run(); + + BOOST_TEST_EQ(n, 5u); + BOOST_TEST_EQ(std::memcmp(buf, msg, 5), 0); + // ...and the first read is what flipped the flag. + BOOST_TEST(::fcntl(fds[0], F_GETFL) & O_NONBLOCK); + + ::close(fds[1]); + } + + void testWaitDoesNotTouchFlags() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[0])); + + char const msg[] = "x"; + BOOST_TEST_EQ(::write(fds[1], msg, 1), 1); + + std::error_code ec; + bool done = false; + auto waiter = [&]() -> capy::task<> { + auto [wec] = co_await d.wait(wait_type::read); + ec = wec; + done = true; + }; + capy::run_async(ioc.get_executor())(waiter()); + ioc.run(); + + BOOST_TEST(done); + BOOST_TEST(!ec); + // wait() alone must never modify a descriptor someone else owns. + BOOST_TEST_EQ(::fcntl(fds[0], F_GETFL) & O_NONBLOCK, 0); + // ...and it consumed nothing. + char buf[4]{}; + BOOST_TEST_EQ(::read(fds[0], buf, 1), 1); + + ::close(fds[1]); + } + + void testWaitReadParksUntilDataArrives() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[0])); + + bool ready = false; + std::error_code ec; + auto waiter = [&]() -> capy::task<> { + auto [wec] = co_await d.wait(wait_type::read); + ec = wec; + ready = true; + }; + auto writer = [&]() -> capy::task<> { + co_await delay(std::chrono::milliseconds(10)); + char const msg[] = "z"; + BOOST_TEST_EQ(::write(fds[1], msg, 1), 1); + }; + + auto ex = ioc.get_executor(); + capy::run_async(ex)(waiter()); + capy::run_async(ex)(writer()); + ioc.run(); + + BOOST_TEST(ready); + BOOST_TEST(!ec); + ::close(fds[1]); + } + + void testWaitWriteOnFullPipe() + { + // Registration latches write_ready on an empty pipe; the probe + // must not let that stale flag report a full pipe as writable. + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[1])); + + // Fill the pipe. The write end must be nonblocking to do this + // without deadlocking, and this is the caller's own fd, so the + // test sets the flag itself rather than relying on the library. + int flags = ::fcntl(fds[1], F_GETFL); + BOOST_TEST_EQ(::fcntl(fds[1], F_SETFL, flags | O_NONBLOCK), 0); + char block[4096]; + std::memset(block, 'a', sizeof(block)); + while (::write(fds[1], block, sizeof(block)) > 0) + { + } + + bool ready = false; + auto waiter = [&]() -> capy::task<> { + auto [wec] = co_await d.wait(wait_type::write); + BOOST_TEST(!wec); + ready = true; + }; + auto drainer = [&]() -> capy::task<> { + co_await delay(std::chrono::milliseconds(10)); + BOOST_TEST_EQ(ready, false); // still parked on the full pipe + char sink[8192]; + BOOST_TEST(::read(fds[0], sink, sizeof(sink)) > 0); + }; + + auto ex = ioc.get_executor(); + capy::run_async(ex)(waiter()); + capy::run_async(ex)(drainer()); + ioc.run(); + + BOOST_TEST(ready); + ::close(fds[0]); + } + + void testWriteSome() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[1])); + + std::size_t n = 0; + auto writer = [&]() -> capy::task<> { + auto [wec, wn] = + co_await d.write_some(capy::const_buffer("abc", 3)); + BOOST_TEST(!wec); + n = wn; + }; + capy::run_async(ioc.get_executor())(writer()); + ioc.run(); + + BOOST_TEST_EQ(n, 3u); + char buf[4]{}; + BOOST_TEST_EQ(::read(fds[0], buf, 3), 3); + BOOST_TEST_EQ(std::memcmp(buf, "abc", 3), 0); + + ::close(fds[0]); + } + + void testCancelPendingRead() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[0])); + + std::error_code ec; + bool done = false; + char buf[8]{}; + auto reader = [&]() -> capy::task<> { + auto [rec, rn] = + co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + ec = rec; + done = true; + (void)rn; + }; + auto canceller = [&]() -> capy::task<> { + d.cancel(); + co_return; + }; + + auto ex = ioc.get_executor(); + capy::run_async(ex)(reader()); + capy::run_async(ex)(canceller()); + ioc.run(); + + BOOST_TEST(done); + BOOST_TEST(ec == capy::cond::canceled); + ::close(fds[1]); + } + + void testReleaseCancelsAndTransfersOwnership() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[0])); + + std::error_code ec; + bool done = false; + char buf[8]{}; + native_handle_type released = -1; + auto reader = [&]() -> capy::task<> { + auto [rec, rn] = + co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + ec = rec; + done = true; + (void)rn; + }; + auto releaser = [&]() -> capy::task<> { + released = d.release(); + co_return; + }; + + auto ex = ioc.get_executor(); + capy::run_async(ex)(reader()); + capy::run_async(ex)(releaser()); + ioc.run(); + + BOOST_TEST(done); + BOOST_TEST(ec == capy::cond::canceled); + BOOST_TEST_EQ(d.is_open(), false); + BOOST_TEST_EQ(released, fds[0]); + // The released fd is still open and ours to close. + BOOST_TEST(::fcntl(released, F_GETFD) >= 0); + ::close(released); + ::close(fds[1]); + } + + void testEofOnClosedWriteEnd() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[0])); + ::close(fds[1]); + + std::error_code ec; + std::size_t n = 1; + char buf[8]{}; + auto reader = [&]() -> capy::task<> { + auto [rec, rn] = + co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + ec = rec; + n = rn; + }; + capy::run_async(ioc.get_executor())(reader()); + ioc.run(); + + BOOST_TEST_EQ(n, 0u); + BOOST_TEST(ec == capy::cond::eof); + } + + void testParkedReadSeesEofWhenWriterCloses() + { + // Distinct from testEofOnClosedWriteEnd, where the write end is + // already closed when read_some() runs and ::read returns 0 on + // the fast path without any backend event. Here the read + // genuinely parks first, so the hangup has to arrive as a + // backend event: epoll reports EPOLLHUP alone for a pipe read + // end -- no EPOLLIN, no EPOLLERR -- and the registration is + // edge-triggered, so a mapping that drops it hangs forever. + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[0])); + + bool done = false; + std::error_code ec; + std::size_t n = 1; + char buf[8]{}; + auto reader = [&]() -> capy::task<> { + auto [rec, rn] = + co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + ec = rec; + n = rn; + done = true; + }; + auto closer = [&]() -> capy::task<> { + co_await delay(std::chrono::milliseconds(10)); + BOOST_TEST_EQ(done, false); // still parked on the empty pipe + ::close(fds[1]); + + // Bound the park so a backend that never delivers the + // hangup fails the assertions below instead of wedging + // ioc.run() for good. + for (int i = 0; i < 50 && !done; ++i) + co_await delay(std::chrono::milliseconds(10)); + if (!done) + d.cancel(); + }; + + auto ex = ioc.get_executor(); + capy::run_async(ex)(reader()); + capy::run_async(ex)(closer()); + ioc.run(); + + BOOST_TEST(done); + BOOST_TEST(ec == capy::cond::eof); + BOOST_TEST_EQ(n, 0u); + } + + void testWaitReadOnHungUpPipeIsReadiness() + { + // A pipe whose writer closed is readable-at-EOF, not faulted: + // the wait-then-read idiom the guide teaches must hand the + // caller a clean wait and let the following read report eof. + // The reactors reach this through poll(POLLIN) returning + // POLLHUP; io_uring through a POLL_ADD revents carrying + // POLLHUP, which must not be turned into a fault. + for (bool close_before_wait : {true, false}) + { + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[0])); + if (close_before_wait) + ::close(fds[1]); + + bool done = false; + std::error_code ec; + auto waiter = [&]() -> capy::task<> { + auto [wec] = co_await d.wait(wait_type::read); + ec = wec; + done = true; + }; + auto closer = [&]() -> capy::task<> { + co_await delay(std::chrono::milliseconds(10)); + if (!close_before_wait) + { + BOOST_TEST_EQ(done, false); // parked, pipe still open + ::close(fds[1]); + } + for (int i = 0; i < 50 && !done; ++i) + co_await delay(std::chrono::milliseconds(10)); + if (!done) + d.cancel(); + }; + + auto ex = ioc.get_executor(); + capy::run_async(ex)(waiter()); + capy::run_async(ex)(closer()); + ioc.run(); + + BOOST_TEST(done); + BOOST_TEST(!ec); + } + } + + void testGatherReadWrite() + { + // Every other case here is single-buffer, which takes the + // ::read / write_one fast path; this is the only coverage of + // the ::readv / ::writev gather forms. + io_context ioc(Backend); + posix_descriptor w(ioc); + posix_descriptor r(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!w.assign(fds[1])); + BOOST_TEST(!r.assign(fds[0])); + + char const head[] = "hello "; + char const tail[] = "world"; + char got_head[6]{}; + char got_tail[5]{}; + std::size_t wn = 0; + std::size_t rn = 0; + + auto body = [&]() -> capy::task<> { + std::array out = { + capy::const_buffer(head, 6), capy::const_buffer(tail, 5)}; + auto [wec, n1] = co_await w.write_some(out); + BOOST_TEST(!wec); + wn = n1; + + std::array in = { + capy::mutable_buffer(got_head, sizeof(got_head)), + capy::mutable_buffer(got_tail, sizeof(got_tail))}; + auto [rec, n2] = co_await r.read_some(in); + BOOST_TEST(!rec); + rn = n2; + }; + capy::run_async(ioc.get_executor())(body()); + ioc.run(); + + BOOST_TEST_EQ(wn, 11u); + BOOST_TEST_EQ(rn, 11u); + BOOST_TEST_EQ(std::memcmp(got_head, "hello ", 6), 0); + BOOST_TEST_EQ(std::memcmp(got_tail, "world", 5), 0); + } + + void testReassignRearmsNonblockingLatch() + { + // A stale nonblocking_ == true would leave the newly adopted + // descriptor blocking while the library believed otherwise. + io_context ioc(Backend); + posix_descriptor d(ioc); + int first[2]; + int second[2]; + make_pipe(first); + make_pipe(second); + BOOST_TEST(!d.assign(first[0])); + + BOOST_TEST_EQ(::write(first[1], "a", 1), 1); + + char buf[8]{}; + auto body = [&]() -> capy::task<> { + auto [ec1, n1] = + co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + BOOST_TEST(!ec1); + BOOST_TEST_EQ(n1, 1u); + BOOST_TEST(::fcntl(first[0], F_GETFL) & O_NONBLOCK); + + // Adopting over the armed descriptor must not inherit its + // latch; the fresh pipe end is still blocking. + BOOST_TEST(!d.assign(second[0])); + BOOST_TEST_EQ(::fcntl(second[0], F_GETFL) & O_NONBLOCK, 0); + + BOOST_TEST_EQ(::write(second[1], "xy", 2), 2); + auto [ec2, n2] = + co_await d.read_some(capy::mutable_buffer(buf, sizeof(buf))); + BOOST_TEST(!ec2); + BOOST_TEST_EQ(n2, 2u); + BOOST_TEST_EQ(std::memcmp(buf, "xy", 2), 0); + BOOST_TEST(::fcntl(second[0], F_GETFL) & O_NONBLOCK); + }; + capy::run_async(ioc.get_executor())(body()); + ioc.run(); + + ::close(first[1]); + ::close(second[1]); + } + + void testParkedWriteOnBrokenPipeReportsEpipe() + { + // The reactor's error probe is getsockopt(SO_ERROR), which fails + // with ENOTSOCK on a pipe. The write must genuinely park before + // the peer closes: a fast-path completion (e.g. closing the read + // end before writing) never reaches invoke_deferred_io's + // ENOTSOCK arm and would pass even with the bug present. + // + // Every suite shares one process, so the disposition has to go + // back the way it was found. + struct sigpipe_guard + { + void (*prev)(int) = ::signal(SIGPIPE, SIG_IGN); + ~sigpipe_guard() + { + ::signal(SIGPIPE, prev); + } + } restore_sigpipe; + + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[1])); + + // Fill the pipe so write_some() genuinely parks. The write end + // is the caller's own fd, so the test sets O_NONBLOCK itself + // rather than relying on the library (see testWaitWriteOnFullPipe). + int flags = ::fcntl(fds[1], F_GETFL); + BOOST_TEST_EQ(::fcntl(fds[1], F_SETFL, flags | O_NONBLOCK), 0); + char block[4096]; + std::memset(block, 'a', sizeof(block)); + while (::write(fds[1], block, sizeof(block)) > 0) + { + } + + bool done = false; + std::error_code ec; + auto writer = [&]() -> capy::task<> { + auto [wec, wn] = co_await d.write_some(capy::const_buffer("x", 1)); + ec = wec; + done = true; + (void)wn; + }; + auto closer = [&]() -> capy::task<> { + co_await delay(std::chrono::milliseconds(10)); + BOOST_TEST_EQ(done, false); // still parked on the full pipe + ::close(fds[0]); + }; + + auto ex = ioc.get_executor(); + capy::run_async(ex)(writer()); + capy::run_async(ex)(closer()); + ioc.run(); + + BOOST_TEST(done); + BOOST_TEST(ec == std::errc::broken_pipe); + BOOST_TEST(ec != std::errc::not_a_socket); + } + + void testWaitErrorParksThenNamesRealCode() + { + // Same ENOTSOCK hazard as the write case above, but through the + // wait_error_op arm: wait(error) must genuinely park (the fd is + // still healthy when wait() probes it) so the later dispatch + // reaches invoke_deferred_io rather than the speculative + // fast-path probe in do_wait(), which never touches SO_ERROR. + // + // Two backends never wake this wait, so on them the bound + // below is what ends it and the error assertions are skipped. + // select's exceptional set does not cover a pipe whose peer + // closed (verified: except_fds never comes back set for it, + // unlike a TCP reset). kqueue raises EV_EOF on both filters for + // the hangup, with fflags == 0 on each (measured on Darwin), and + // kqueue_scheduler raises reactor_event_error only when + // fflags != 0; a TCP reset does set fflags, which is why + // tcp_socket.cpp's + // testWaitForErrorThenWait -- the precedent this copies -- + // needs no kqueue exemption. + io_context ioc(Backend); + posix_descriptor d(ioc); + int fds[2]; + make_pipe(fds); + BOOST_TEST(!d.assign(fds[1])); + + bool done = false; + std::error_code ec; + auto waiter = [&]() -> capy::task<> { + auto [wec] = co_await d.wait(wait_type::error); + ec = wec; + done = true; + }; + auto closer = [&]() -> capy::task<> { + co_await delay(std::chrono::milliseconds(10)); + BOOST_TEST_EQ(done, false); // still parked, no error yet + ::close(fds[0]); + + // Bound the wait: a backend that never reports this as an + // exceptional condition would otherwise park it for good. + for (int i = 0; i < 20 && !done; ++i) + co_await delay(std::chrono::milliseconds(10)); + if (!done) + d.cancel(); + }; + + auto ex = ioc.get_executor(); + capy::run_async(ex)(waiter()); + capy::run_async(ex)(closer()); + ioc.run(); + + BOOST_TEST(done); +#if BOOST_COROSIO_HAS_SELECT + constexpr bool is_select = + std::is_same_v, select_t>; +#else + constexpr bool is_select = false; +#endif +#if BOOST_COROSIO_HAS_KQUEUE + constexpr bool is_kqueue = + std::is_same_v, kqueue_t>; +#else + constexpr bool is_kqueue = false; +#endif + if constexpr (!is_select && !is_kqueue) + { + // Never the cancel: that would say the close reached + // nothing and the bound is what ended the wait. + BOOST_TEST(ec != capy::cond::canceled); + BOOST_TEST(ec == std::errc::io_error); + } + } + + void testWaitOnClosedDescriptor() + { + io_context ioc(Backend); + posix_descriptor d(ioc); + + std::error_code ec; + bool done = false; + auto waiter = [&]() -> capy::task<> { + auto [wec] = co_await d.wait(wait_type::read); + ec = wec; + done = true; + }; + capy::run_async(ioc.get_executor())(waiter()); + ioc.run(); + + BOOST_TEST(done); + BOOST_TEST(ec == std::errc::bad_file_descriptor); + } + + void run() + { + testConstruction(); + testAssignRejectsRegularFile(); + testAssignRejectsBadFd(); + testAssignRejectsSelf(); + testFailedAssignLeavesPriorStateIntact(); + testAssignDoesNotTouchFlagsButReadDoes(); + testWaitDoesNotTouchFlags(); + testWaitReadParksUntilDataArrives(); + testWaitWriteOnFullPipe(); + testWriteSome(); + testCancelPendingRead(); + testReleaseCancelsAndTransfersOwnership(); + testEofOnClosedWriteEnd(); + testParkedReadSeesEofWhenWriterCloses(); + testWaitReadOnHungUpPipeIsReadiness(); + testGatherReadWrite(); + testReassignRearmsNonblockingLatch(); + testParkedWriteOnBrokenPipeReportsEpipe(); + testWaitErrorParksThenNamesRealCode(); + testWaitOnClosedDescriptor(); + } +}; + +COROSIO_BACKEND_TESTS(posix_descriptor_test, "boost.corosio.posix_descriptor") + +} // namespace boost::corosio + +#endif // BOOST_COROSIO_POSIX diff --git a/test/unit/random_access_file.cpp b/test/unit/random_access_file.cpp index 8186aa653..22f47a105 100644 --- a/test/unit/random_access_file.cpp +++ b/test/unit/random_access_file.cpp @@ -46,6 +46,7 @@ #include "temp_path.hpp" #if BOOST_COROSIO_POSIX +#include #include #include #else @@ -719,6 +720,76 @@ struct random_access_file_test BOOST_TEST(!resumed); } +#if BOOST_COROSIO_POSIX + // assign() calls cancel() before close_file() (POSIX/uring only): + // raf_op::do_work reads fd_ at pool-execution time, not at post + // time, so without the cancel a read_at queued before assign() + // would silently complete against the newly adopted file instead + // of being cancelled. + void testAssignCancelsInFlightRead() + { +#if BOOST_COROSIO_HAS_URING + // uring's close_file() cancels internally via + // sched_->cancel_and_flush(fd_); this pins the POSIX pool path, + // where the cancel used to live in the service, not assign(). + if constexpr ( + std::is_same_v, uring_t>) + return; +#endif + temp_file tmp1("raf_assign_cancel_a_", "OLDOLDOLD"); + temp_file tmp2("raf_assign_cancel_b_", "NEWNEWNEW"); + + // blocker must outlive ioc: the pool joins its workers while + // the context is being destroyed, and pool_release_gate's + // shutdown() calls blocker.release() at that point (see + // pool_teardown.hpp). Declaring it after ioc destroys it first, + // and ioc's destructor then releases an already-dead object. + test::pool_blocker blocker; + io_context ioc(Backend); + BOOST_TEST(test::park_pool_worker(ioc, blocker)); + + random_access_file f(ioc); + BOOST_TEST(!f.open(tmp1.path, file_base::read_only)); + + bool resumed = false; + std::error_code result_ec = {}; + std::size_t result_bytes = 0; + char buf[16] = {}; + + auto reader = [&]() -> capy::task<> { + auto [ec, n] = co_await f.read_some_at( + 0, capy::mutable_buffer(buf, sizeof(buf))); + result_ec = ec; + result_bytes = n; + resumed = true; + }; + + std::optional ex; + std::optional env; + std::optional> parked; + ex.emplace(ioc.get_executor()); + env.emplace(capy::io_env{*ex, std::stop_token{}, nullptr}); + parked.emplace(reader()); + // Synchronously runs the coroutine up to its first suspension, + // which posts the read to the pool -- queued behind the parked + // worker, not yet executed. + parked->await_suspend(std::noop_coroutine(), &*env).resume(); + + int fd2 = ::open(tmp2.path.c_str(), O_RDONLY); + BOOST_TEST(fd2 >= 0); + BOOST_TEST(!f.assign(static_cast(fd2))); + + blocker.release(); + ioc.run(); + + BOOST_TEST(resumed); + BOOST_TEST(result_ec == capy::cond::canceled); + BOOST_TEST_EQ(result_bytes, 0u); + // Must not have completed against the newly adopted file's data. + BOOST_TEST(std::memcmp(buf, "NEWNEWNEW", 9) != 0); + } +#endif + void run() { testConstruction(); @@ -758,8 +829,10 @@ struct random_access_file_test testClosedAtOpsComplete(); testStopRaceReportsTransfer(); #if BOOST_COROSIO_POSIX - testSyncOnPipeFails(); + testAssignPipeRejected(); testHugeOffsetFails(); + testFailedAssignLeavesFileOpen(); + testSelfAssignRejected(); #endif testWrongDirectionIoFails(); testResizeReadOnlyFails(); @@ -777,6 +850,7 @@ struct random_access_file_test // POSIX file work runs on the pool; IOCP uses overlapped I/O. testDestroyWithPoolWorkQueued(); testReadWriteAtAfterPoolShutdown(); + testAssignCancelsInFlightRead(); #endif #if !COROSIO_TEST_HAS_ASAN @@ -834,6 +908,9 @@ struct random_access_file_test random_access_file f(ioc); BOOST_TEST(!f.open(tmp1.path, file_base::read_only)); + // Only read back under the POSIX guard below; IOCP's assign() + // is close-first and does not offer the property being pinned. + [[maybe_unused]] auto held = f.native_handle(); #if BOOST_COROSIO_HAS_IOCP HANDLE h = ::CreateFileW( @@ -851,21 +928,66 @@ struct random_access_file_test BOOST_TEST(!f.assign(raw)); BOOST_TEST(f.is_open()); BOOST_TEST_EQ(f.size(), 6u); + +#if BOOST_COROSIO_POSIX + // Pin the property the (now-removed) wrapper-level close() used + // to provide: a *successful* assign still closes the fd it + // replaced, not just the one it rejects. + errno = 0; + BOOST_TEST(::fcntl(held, F_GETFD) < 0); + BOOST_TEST_EQ(errno, EBADF); +#endif } #if BOOST_COROSIO_POSIX - void testSyncOnPipeFails() + // These validate assign()'s new fail-before-mutate contract, which so + // far is POSIX/uring-only -- IOCP's file services are unified (and + // gated) in a later stage. + void testFailedAssignLeavesFileOpen() + { + io_context ioc(Backend); + random_access_file f(ioc); + + temp_file tmp("raf_assign_reject_"); + BOOST_TEST( + !f.open(tmp.path, file_base::read_write | file_base::create)); + auto held = f.native_handle(); + + // A rejected assign must not close the file we already hold. + BOOST_TEST(f.assign(-1) == std::errc::bad_file_descriptor); + BOOST_TEST_EQ(f.is_open(), true); + BOOST_TEST_EQ(f.native_handle(), held); + } + + void testSelfAssignRejected() + { + io_context ioc(Backend); + random_access_file f(ioc); + + temp_file tmp("raf_assign_self_"); + BOOST_TEST( + !f.open(tmp.path, file_base::read_write | file_base::create)); + + BOOST_TEST(f.assign(f.native_handle()) == std::errc::invalid_argument); + BOOST_TEST_EQ(f.is_open(), true); + } + + void testAssignPipeRejected() { + // A pipe has no file position, so validate_file_fd now rejects + // it at assign() -- before this test relied on fsync/fdatasync + // surfacing a runtime error on an adopted pipe fd. io_context ioc(Backend); random_access_file f(ioc); int fds[2]; BOOST_TEST(::pipe(fds) == 0); - BOOST_TEST(!f.assign(static_cast(fds[1]))); - BOOST_TEST(f.sync_data()); - BOOST_TEST(f.sync_all()); - f.close(); + BOOST_TEST( + f.assign(static_cast(fds[1])) == + std::errc::operation_not_supported); + BOOST_TEST(!f.is_open()); ::close(fds[0]); + ::close(fds[1]); } #endif diff --git a/test/unit/stream_file.cpp b/test/unit/stream_file.cpp index 11d3f0204..2466e44a5 100644 --- a/test/unit/stream_file.cpp +++ b/test/unit/stream_file.cpp @@ -46,6 +46,7 @@ #include "temp_path.hpp" #if BOOST_COROSIO_POSIX +#include #include #include #else @@ -749,6 +750,9 @@ struct stream_file_test stream_file f(ioc); BOOST_TEST(!f.open(tmp1.path, file_base::read_only)); + // Only read back under the POSIX guard below; IOCP's assign() + // is close-first and does not offer the property being pinned. + [[maybe_unused]] auto held = f.native_handle(); #if BOOST_COROSIO_HAS_IOCP HANDLE h = ::CreateFileW( @@ -768,23 +772,66 @@ struct stream_file_test BOOST_TEST(!f.assign(raw)); BOOST_TEST(f.is_open()); BOOST_TEST_EQ(f.size(), 6u); + +#if BOOST_COROSIO_POSIX + // Pin the property the (now-removed) wrapper-level close() used + // to provide: a *successful* assign still closes the fd it + // replaced, not just the one it rejects. + errno = 0; + BOOST_TEST(::fcntl(held, F_GETFD) < 0); + BOOST_TEST_EQ(errno, EBADF); +#endif } #if BOOST_COROSIO_POSIX - void testSyncOnPipeFails() + // These validate assign()'s new fail-before-mutate contract, which so + // far is POSIX/uring-only -- IOCP's file services are unified (and + // gated) in a later stage. + void testFailedAssignLeavesFileOpen() + { + io_context ioc(Backend); + stream_file f(ioc); + + temp_file tmp("sf_assign_reject_"); + BOOST_TEST( + !f.open(tmp.path, file_base::read_write | file_base::create)); + auto held = f.native_handle(); + + // A rejected assign must not close the file we already hold. + BOOST_TEST(f.assign(-1) == std::errc::bad_file_descriptor); + BOOST_TEST_EQ(f.is_open(), true); + BOOST_TEST_EQ(f.native_handle(), held); + } + + void testSelfAssignRejected() + { + io_context ioc(Backend); + stream_file f(ioc); + + temp_file tmp("sf_assign_self_"); + BOOST_TEST( + !f.open(tmp.path, file_base::read_write | file_base::create)); + + BOOST_TEST(f.assign(f.native_handle()) == std::errc::invalid_argument); + BOOST_TEST_EQ(f.is_open(), true); + } + + void testAssignPipeRejected() { - // fsync/fdatasync on a pipe reports a genuine runtime error - // through the returned code. + // A pipe has no file position, so validate_file_fd now rejects + // it at assign() -- before this test relied on fsync/fdatasync + // surfacing a runtime error on an adopted pipe fd. io_context ioc(Backend); stream_file f(ioc); int fds[2]; BOOST_TEST(::pipe(fds) == 0); - BOOST_TEST(!f.assign(static_cast(fds[1]))); - BOOST_TEST(f.sync_data()); - BOOST_TEST(f.sync_all()); - f.close(); + BOOST_TEST( + f.assign(static_cast(fds[1])) == + std::errc::operation_not_supported); + BOOST_TEST(!f.is_open()); ::close(fds[0]); + ::close(fds[1]); } #endif @@ -1027,6 +1074,76 @@ struct stream_file_test BOOST_TEST(!resumed); } +#if BOOST_COROSIO_POSIX + // assign() calls cancel() before close_file() (POSIX/uring only): + // do_read_work reads fd_/offset_ at pool-execution time, not at + // post time, so without the cancel a read queued before assign() + // would silently complete against the newly adopted file instead + // of being cancelled. + void testAssignCancelsInFlightRead() + { +#if BOOST_COROSIO_HAS_URING + // uring's close_file() cancels internally via + // sched_->cancel_and_flush(fd_); this pins the POSIX pool path, + // where the cancel used to live in the service, not assign(). + if constexpr ( + std::is_same_v, uring_t>) + return; +#endif + temp_file tmp1("sf_assign_cancel_a_", "OLDOLDOLD"); + temp_file tmp2("sf_assign_cancel_b_", "NEWNEWNEW"); + + // blocker must outlive ioc: the pool joins its workers while + // the context is being destroyed, and pool_release_gate's + // shutdown() calls blocker.release() at that point (see + // pool_teardown.hpp). Declaring it after ioc destroys it first, + // and ioc's destructor then releases an already-dead object. + test::pool_blocker blocker; + io_context ioc(Backend); + BOOST_TEST(test::park_pool_worker(ioc, blocker)); + + stream_file f(ioc); + BOOST_TEST(!f.open(tmp1.path, file_base::read_only)); + + bool resumed = false; + std::error_code result_ec = {}; + std::size_t result_bytes = 0; + char buf[16] = {}; + + auto reader = [&]() -> capy::task<> { + auto [ec, n] = + co_await f.read_some(capy::mutable_buffer(buf, sizeof(buf))); + result_ec = ec; + result_bytes = n; + resumed = true; + }; + + std::optional ex; + std::optional env; + std::optional> parked; + ex.emplace(ioc.get_executor()); + env.emplace(capy::io_env{*ex, std::stop_token{}, nullptr}); + parked.emplace(reader()); + // Synchronously runs the coroutine up to its first suspension, + // which posts the read to the pool -- queued behind the parked + // worker, not yet executed. + parked->await_suspend(std::noop_coroutine(), &*env).resume(); + + int fd2 = ::open(tmp2.path.c_str(), O_RDONLY); + BOOST_TEST(fd2 >= 0); + BOOST_TEST(!f.assign(static_cast(fd2))); + + blocker.release(); + ioc.run(); + + BOOST_TEST(resumed); + BOOST_TEST(result_ec == capy::cond::canceled); + BOOST_TEST_EQ(result_bytes, 0u); + // Must not have completed against the newly adopted file's data. + BOOST_TEST(std::memcmp(buf, "NEWNEWNEW", 9) != 0); + } +#endif + void run() { testConstruction(); @@ -1069,10 +1186,14 @@ struct stream_file_test testClosedFileErrors(); testResizeReadOnlyFails(); #if BOOST_COROSIO_POSIX - testSyncOnPipeFails(); + testAssignPipeRejected(); #endif testWrongDirectionIoFails(); testAssignOverOpenAdopts(); +#if BOOST_COROSIO_POSIX + testFailedAssignLeavesFileOpen(); + testSelfAssignRejected(); +#endif testSeekNegative(); testCancelWithStoppedToken(); testStopRaceReportsTransfer(); @@ -1081,6 +1202,7 @@ struct stream_file_test // POSIX file work runs on the pool; IOCP uses overlapped I/O. testDestroyWithPoolWorkQueued(); testReadWriteAfterPoolShutdown(); + testAssignCancelsInFlightRead(); #endif #if !COROSIO_TEST_HAS_ASAN diff --git a/test/unit/teardown_inflight.cpp b/test/unit/teardown_inflight.cpp index a11978ac7..66462427d 100644 --- a/test/unit/teardown_inflight.cpp +++ b/test/unit/teardown_inflight.cpp @@ -22,8 +22,8 @@ #include #include #include +#include #include -#include #include #include #include @@ -223,7 +223,7 @@ struct uring_teardown_test BOOST_TEST(!resumed); } - void testDestroyWithPendingFileOps() + void testDestroyWithPendingDescriptorOps() { // The counterpart end of each pipe stays raw and open past the // context so the drained ops never hit a broken pipe. @@ -236,16 +236,16 @@ struct uring_teardown_test { io_context ioc(uring); auto reader = [&]() -> capy::task<> { - stream_file f(ioc); - std::ignore = f.assign(static_cast(rp[0])); + posix_descriptor f(ioc); + BOOST_TEST(!f.assign(static_cast(rp[0]))); char buf[16]; std::ignore = co_await f.read_some( capy::mutable_buffer(buf, sizeof(buf))); read_resumed = true; }; auto writer = [&]() -> capy::task<> { - stream_file f(ioc); - std::ignore = f.assign(static_cast(wp[1])); + posix_descriptor f(ioc); + BOOST_TEST(!f.assign(static_cast(wp[1]))); char big[4096] = {}; std::ignore = co_await f.write_some(capy::const_buffer(big, sizeof(big))); @@ -301,12 +301,23 @@ struct uring_teardown_test // undispatched when the context dies, so the drain must run its // handler ownerless. - void testDestroyWithBrokenPipeWriteSurvives() + void testDestroyDrainsBothEndsOfOnePipe() { - // Both ends of one pipe are wrapped, with a write parked on - // the full pipe. Service shutdown closes the read end first, - // so the flushed write executes against a broken pipe; the - // library must absorb the SIGPIPE instead of dying. + // Both ends of one pipe are wrapped, with a read parked on the + // empty end and a write parked on the full one, so service + // shutdown closes one end while the other's op is still in + // flight. Neither coroutine may resume. + // + // This was testDestroyWithBrokenPipeWriteSurvives, and it no + // longer earns that name. While the ends were stream_files the + // flushed write ran against a broken pipe and the SIGPIPE had + // to be absorbed -- deleting uring_stream_file::close_file()'s + // scoped_sigpipe_block kills the process. uring_descriptor + // transfers are two-phase, so after two run_one() calls what + // is in the ring is a poll_add, which flushes harmlessly; + // deleting uring_descriptor::close_descriptor()'s guard leaves + // this passing. The SIGPIPE coverage is gone, not relocated. + // What survives is the drain coverage named above. int p[2]; BOOST_TEST(::pipe2(p, O_NONBLOCK) == 0); fill_pipe(p[1]); @@ -315,16 +326,16 @@ struct uring_teardown_test { io_context ioc(uring); auto reader = [&]() -> capy::task<> { - stream_file f(ioc); - std::ignore = f.assign(static_cast(p[0])); + posix_descriptor f(ioc); + BOOST_TEST(!f.assign(static_cast(p[0]))); char buf[16]; std::ignore = co_await f.read_some( capy::mutable_buffer(buf, sizeof(buf))); read_resumed = true; }; auto writer = [&]() -> capy::task<> { - stream_file f(ioc); - std::ignore = f.assign(static_cast(p[1])); + posix_descriptor f(ioc); + BOOST_TEST(!f.assign(static_cast(p[1]))); char big[4096] = {}; std::ignore = co_await f.write_some(capy::const_buffer(big, sizeof(big))); @@ -413,7 +424,7 @@ struct uring_teardown_test BOOST_TEST_LT(resumed, 2); } - void testDestroyWithQueuedFileOps() + void testDestroyWithQueuedDescriptorOps() { int rp[2], wp[2]; BOOST_TEST(::pipe2(rp, O_NONBLOCK) == 0); @@ -425,8 +436,8 @@ struct uring_teardown_test io_context ioc(uring); auto reader = [](io_context& ctx, int fd, int& count) -> capy::task<> { - stream_file f(ctx); - std::ignore = f.assign(static_cast(fd)); + posix_descriptor f(ctx); + BOOST_TEST(!f.assign(static_cast(fd))); char buf[4]; std::ignore = co_await f.read_some( capy::mutable_buffer(buf, sizeof(buf))); @@ -437,8 +448,8 @@ struct uring_teardown_test }; auto writer = [](io_context& ctx, int fd, int& count) -> capy::task<> { - stream_file f(ctx); - std::ignore = f.assign(static_cast(fd)); + posix_descriptor f(ctx); + BOOST_TEST(!f.assign(static_cast(fd))); std::ignore = co_await f.write_some(capy::const_buffer("a", 1)); ++count; std::ignore = co_await f.write_some(capy::const_buffer("b", 1)); @@ -524,12 +535,12 @@ struct uring_teardown_test #if !COROSIO_TEST_HAS_ASAN testDestroyWithPendingSocketWrite(); testDestroyWithPendingDatagramSend(); - testDestroyWithPendingFileOps(); + testDestroyWithPendingDescriptorOps(); testDestroyWithSubmittedRandomAccessOps(); - testDestroyWithBrokenPipeWriteSurvives(); + testDestroyDrainsBothEndsOfOnePipe(); testDestroyWithQueuedSocketWrites(); testDestroyWithQueuedDatagramSends(); - testDestroyWithQueuedFileOps(); + testDestroyWithQueuedDescriptorOps(); testDestroyWithQueuedRandomAccessOps(); testDestroyWithParkedLocalAccept(); #endif