Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 8 additions & 2 deletions .github/workflows/kernel-matrix.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,13 +7,19 @@ name: kernel-matrix
# QEMU/KVM on the runner. Each job writes a detail table to its step summary and
# uploads its result; the final `matrix` job pivots them into one ✅/❌ grid.
#
# apiwatch links one object per bpf/<name>/ directory: socket (fentry on
# tcp_sendmsg/tcp_recvmsg) and the TLS taps (ssl, ssl_ex, rustls), because a
# apiwatch links one object per bpf/<name>/ directory: socket (kprobes on
# tcp_sendmsg/tcp_recvmsg and a kretprobe on tcp_recvmsg) and the TLS taps
# (ssl, ssl_ex, rustls: uprobes on the library, plus kprobes on
# tcp_sendmsg/tcp_recvmsg to find each connection's socket), because a
# uprobe tap must load independently of the kernel-global probes. The
# matrix therefore has a row per (object, program): peer_sendmsg and
# peer_recvmsg recur in every TLS tap, so a bare program name would not be
# unique. The objects and their loaders come from yeet-src/httpscope.
#
# The kernel hooks are kprobes rather than fentry because fentry cannot
# attach on arm64 before 6.4. veristat loads programs but does not attach
# them, so this matrix would not catch an attach failure like that one.
#
# Tune `matrix.kernel` to the kernel lines your script must support (`6.6`,
# `bpf-next`, …); available lines live at
# https://quay.io/repository/lvh-images/kind?tab=tags. Each line is resolved to
Expand Down
395 changes: 395 additions & 0 deletions README.md

Large diffs are not rendered by default.

19 changes: 11 additions & 8 deletions bpf/include/tap.h
Original file line number Diff line number Diff line change
Expand Up @@ -53,7 +53,7 @@ struct {
* thread about to call tcp_sendmsg on that connection's socket (for a
* socket BIO, inside the same call; for a memory BIO, on the same tick),
* and the thread inside SSL_read is the one calling tcp_recvmsg. So the
* tap notes (thread → conn) on the way in, and an fentry on the two
* tap notes (thread → conn) on the way in, and a kprobe on the two
* socket calls turns the next hit on that thread into a (conn → socket)
* binding, emitted once per connection.
*
Expand Down Expand Up @@ -108,18 +108,21 @@ static __always_inline void peer_hit(struct sock *sk)
bpf_ringbuf_submit(e, 0);
}

/* Kernel-global, BTF-typed, and auto-attached with the object. Only the
* first argument is read, so the declaration holds across kernels that
* changed the rest of the signature. */
SEC("fentry/tcp_sendmsg")
int BPF_PROG(peer_sendmsg, struct sock *sk)
/* Kernel-global and auto-attached with the object. Kprobes and not
* fentry, because fentry cannot attach on arm64 before 6.4 (Graviton on
* Amazon Linux 2023, for one) and kprobes attach everywhere this runs;
* `sk` is then a bare register, read only through BPF_CORE_READ in
* read_flow(). Only the first argument is read, so the declaration holds
* across kernels that changed the rest of the signature. */
SEC("kprobe/tcp_sendmsg")
int BPF_KPROBE(peer_sendmsg, struct sock *sk)
{
peer_hit(sk);
return 0;
}

SEC("fentry/tcp_recvmsg")
int BPF_PROG(peer_recvmsg, struct sock *sk)
SEC("kprobe/tcp_recvmsg")
int BPF_KPROBE(peer_recvmsg, struct sock *sk)
{
peer_hit(sk);
return 0;
Expand Down
51 changes: 33 additions & 18 deletions bpf/socket/socket.bpf.c
Original file line number Diff line number Diff line change
@@ -1,11 +1,17 @@
/* The socket tap: plaintext capture at tcp_sendmsg / tcp_recvmsg for the
* connections you point it at.
*
* Everything here is a kernel-global fentry/fexit, so this object always
* loads — there is no uprobe in it to fail an attach. That is why it is
* kept apart from the TLS taps (start() rejects an object with an
* unattached uprobe): a process with no OpenSSL must not be able to take
* plain-HTTP capture down with it.
* Everything here is a kernel-global kprobe/kretprobe, so there is no
* uprobe in it to fail an attach. That is why it is kept apart from the
* TLS taps (start() rejects an object with an unattached uprobe): a
* process with no OpenSSL must not be able to take plain-HTTP capture
* down with it.
*
* Kprobes and not fentry/fexit, because fentry cannot attach on arm64
* before 6.4 (no ftrace direct calls there; Graviton on Amazon Linux 2023
* is one such machine), and kprobes attach everywhere this runs. The
* arguments are then bare registers rather than BTF pointers, so every
* read through them is a BPF_CORE_READ or a probe read.
*
* tcp_sendmsg carries what a process sends, in the user iovec, at entry;
* tcp_recvmsg's buffer is filled by return, so the entry stashes the
Expand Down Expand Up @@ -89,8 +95,11 @@ struct {
__uint(max_entries, 1 << 24);
} frames SEC(".maps");

/* A read's entry, kept per thread until its return. LRU because a
* kretprobe can miss a return (see on_recvmsg_exit), and a slot whose
* return never ran must not hold the map forever. */
struct {
__uint(type, BPF_MAP_TYPE_HASH);
__uint(type, BPF_MAP_TYPE_LRU_HASH);
__type(key, __u64);
__type(value, struct read_args);
__uint(max_entries, 10240);
Expand Down Expand Up @@ -179,8 +188,8 @@ static __always_inline int iter_first(struct msghdr *msg, __u64 *base, __u64 *le
* iovec segments (a header block and a body written with one writev),
* so each segment is emitted on its own and the decoder concatenates
* per connection. */
SEC("fentry/tcp_sendmsg")
int BPF_PROG(on_sendmsg, struct sock *sk, struct msghdr *msg, size_t size)
SEC("kprobe/tcp_sendmsg")
int BPF_KPROBE(on_sendmsg, struct sock *sk, struct msghdr *msg, size_t size)
{
if ((long) size <= 0)
return 0;
Expand Down Expand Up @@ -217,20 +226,24 @@ int BPF_PROG(on_sendmsg, struct sock *sk, struct msghdr *msg, size_t size)

/* int tcp_recvmsg(struct sock *sk, struct msghdr *msg, size_t len, ...):
* the destination is filled by return. Stash (sk, where the iterator
* will write, how much room) at entry; the exit reads the return count
* with the helper — the arity changed in 5.19 — and copies min(ret,
* room). `flags` rides along so the record can say a read was a peek.
* will write, how much room) at entry; a return probe has no arguments
* to read, so the exit takes all of that from the stash, reads the count
* from the return register, and copies min(ret, room). `flags` rides
* along so the record can say a read was a peek.
*
* Two things that were wrong here once: the iterator's `iov_offset` was
* ignored, so a second tcp_recvmsg round inside one syscall re-read the
* start of the buffer; and the int return was taken zero-extended, so
* -EAGAIN became a huge count capped to the room — a record full of
* whatever the buffer held before. Both showed up as a response head
* where curl's body should have been. */
SEC("fentry/tcp_recvmsg")
int BPF_PROG(on_recvmsg_enter, struct sock *sk, struct msghdr *msg, size_t len, int flags)
SEC("kprobe/tcp_recvmsg")
int BPF_KPROBE(on_recvmsg_enter, struct sock *sk, struct msghdr *msg, size_t len, int flags)
{
__u64 id = bpf_get_current_pid_tgid();
/* A return the kretprobe missed never cleared its slot. Clear it now,
* so this call's return cannot copy from the previous call's buffer. */
bpf_map_delete_elem(&active_reads, &id);
if (!wanted(sk, id >> 32))
return 0;
struct read_args a = { .conn = (__u64) sk, .flags = (__u64) (__u32) flags };
Expand All @@ -240,19 +253,21 @@ int BPF_PROG(on_recvmsg_enter, struct sock *sk, struct msghdr *msg, size_t len,
return 0;
}

SEC("fexit/tcp_recvmsg")
int BPF_PROG(on_recvmsg_exit, struct sock *sk)
/* A kretprobe has a fixed pool of in-flight instances, sized from the
* CPU count, so when more threads than that sit in tcp_recvmsg at once
* some returns are missed and those reads are not captured. */
SEC("kretprobe/tcp_recvmsg")
int BPF_KRETPROBE(on_recvmsg_exit)
{
__u64 id = bpf_get_current_pid_tgid();
struct read_args *a = bpf_map_lookup_elem(&active_reads, &id);
if (!a)
return 0;
struct sock *sk = (struct sock *) a->conn;
__u64 buf = a->buf, cap = a->nread, flags = a->flags;
bpf_map_delete_elem(&active_reads, &id);

__u64 ret = 0;
if (bpf_get_func_ret(ctx, &ret))
return 0;
__u64 ret = PT_REGS_RC(ctx);
/* The int return arrives zero-extended: -EAGAIN is 0xfffffff5 here,
* a very large count, unless it is taken as the int it was. */
long n = (int) ret;
Expand Down
6 changes: 4 additions & 2 deletions build/verify-kernel.sh
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,10 @@
#
# sh build/verify-kernel.sh [bpf-object ...] (default: bin/*.bpf.o)
#
# apiwatch links one object per bpf/<name>/ directory (socket, ssl, ssl_ex,
# rustls), so the default is the whole set:
# apiwatch links one object per bpf/<name>/ directory: socket (kprobes and
# a kretprobe on tcp_sendmsg/tcp_recvmsg) and ssl, ssl_ex, rustls (uprobes,
# plus kprobes on the same two socket calls), so the default is the whole
# set:
# veristat takes several objects in one run and names the file in each row.
#
# Set OUT_CSV=<path> to also write a machine-readable result (file,prog,verdict,
Expand Down
4 changes: 3 additions & 1 deletion src/lib/alerts.js
Original file line number Diff line number Diff line change
Expand Up @@ -132,7 +132,9 @@ export class Alerts {
const body = api.kind === "served"
? `*${row.name}* (port ${api.port} on ${this.host}) answered ${of} requests with a 5xx in the last ${this.window}s.`
: `*${row.name}* returned a 5xx to ${of} calls from ${row.process ?? this.host} in the last ${this.window}s.`;
const related = this.related(api.key);
/* Only for an API this box serves: a stopped local port explains a
* proxy's 502, and says nothing about a third party's 503. */
const related = api.kind === "served" ? this.related(api.key) : null;
return {
title: `${row.name} is returning ${codes.length ? codes.join(" and ") : "5xx"}`,
body: [body, lines.length ? `Latest: ${lines.join(", ")}` : null, related].filter(Boolean).join("\n"),
Expand Down
2 changes: 1 addition & 1 deletion src/lib/capture.js
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
*
* Three sources, one decoder:
*
* socket tap fentry on tcp_sendmsg/tcp_recvmsg (bin/socket.bpf.o):
* socket tap kprobes on tcp_sendmsg/tcp_recvmsg (bin/socket.bpf.o):
* plaintext HTTP/1.x and h2c on any port, both ends of a
* loopback hop, plus the ClientHello of every TLS call
* TLS taps uprobes on SSL_read/SSL_write (+ _ex) and rustls, attached
Expand Down
Loading