Latest Results
feat(node): enforce per-tenant storage quotas and rate limits (#53)
Two limits per tenant: stored bytes, and operations per second. Both
default to unlimited, because a node that started enforcing a limit
nobody configured would reject writes for reasons its operator never
chose.
The unit of a quota is the caller's key plus the caller's value. Not the
encoded key: the tenant-name length prefix and the subspace byte are the
node's framing, and billing a tenant for the length of its own name would
be indefensible. Not the engine's on-disk footprint either, because that
moves with compaction and a quota that drifts under the tenant's feet is
not a quota.
Usage is tracked in memory, because a quota check on the write path
cannot afford a scan, and materialised into the tenant's stats subspace
so an operator or a billing job can read it without asking the node.
Startup recomputes it by scanning the real data rather than trusting the
stored total. That costs nothing — the TTL index needs the same scan, and
the two now share one pass — and it makes the counter self-healing: drift
from a crash between the data write and the counter write, or from a bug
in the accounting, is corrected at the next boot instead of compounding.
A write reads the current entry to charge the difference rather than the
whole value. Counting only additions would inflate usage on every
overwrite, and a quota that wanders upwards without the data growing
rejects legitimate writes for a reason the tenant cannot see. The check
runs before the WAL append, so a refusal cannot leave a record replay
would re-apply.
A write that frees bytes is always admitted, even for a tenant already
over its limit. Otherwise the only way out of an exceeded quota would be
an operator raising it, and a tenant could not fix its own problem.
Rate limiting is admission control, so it runs before any storage access
and covers reads as well as writes: a tenant can saturate a node with
scan just as effectively as with put. A batch costs one operation per op
it contains, since charging per RPC would let any client willing to batch
bypass the limit entirely. A request costing more than the whole
per-second allowance is InvalidArgument rather than ResourceExhausted —
no wait admits it, so reporting it as a rate limit would send the client
into a retry loop that cannot succeed.
governor's keyed limiter would allocate a key per request just to look
one up, on the hottest path in the node. The map is ours, keyed by
Arc<str> so it can be probed with a bare &str; the limiters are
governor's.
Admin API on the existing admin server rather than as gRPC RPCs. That
port is already the operator surface and binds to loopback by default,
whereas an administrative RPC on the tenant-facing API would need an
admin RBAC role that does not exist yet.
Closes #26 refactor(node): derive the delete found flag from the delete itself
The handler read the key and then deleted it as two separate storage
operations, so the flag and the removal described two different points in
time. Since quota accounting already reads the entry, the read that
produced the flag was also redundant.
`NodeService::delete` now returns whether the key held a value, taken
from the read the accounting needs. One read means one point in time: the
flag and the removal can no longer disagree about the same delete.
The flag is still best-effort and this does not change that. Two clients
deleting the same key concurrently can still both be told true. Making it
exact needs an atomic remove-and-return: MemStorage has one for free
since DashMap::remove hands back the old value, fjall does not short of a
transaction, and implementing it for one backend would make the semantics
differ by backend — worse than a documented approximation.
So it is documented where a caller will actually meet it: on the proto
message, on the service method, and in the README. All three say the same
thing — safe to log or display, not safe to branch on, and specifically
not as a lock, a once-only trigger, or a claim on a work item.
While separating the existence check from the accounting: the accounting
plan returns None both for an absent key and for a key outside the data
subspace, so deriving the flag from it would have reported that deleting
a metadata key removed nothing. Split, with a test.
Refs #12 refactor(node): derive the delete found flag from the delete itself
The handler read the key and then deleted it as two separate storage
operations, so the flag and the removal described two different points in
time. Since quota accounting already reads the entry, the read that
produced the flag was also redundant.
`NodeService::delete` now returns whether the key held a value, taken
from the read the accounting needs. One read means one point in time: the
flag and the removal can no longer disagree about the same delete.
The flag is still best-effort and this does not change that. Two clients
deleting the same key concurrently can still both be told true. Making it
exact needs an atomic remove-and-return: MemStorage has one for free
since DashMap::remove hands back the old value, fjall does not short of a
transaction, and implementing it for one backend would make the semantics
differ by backend — worse than a documented approximation.
So it is documented where a caller will actually meet it: on the proto
message, on the service method, and in the README. All three say the same
thing — safe to log or display, not safe to branch on, and specifically
not as a lock, a once-only trigger, or a claim on a work item.
While separating the existence check from the accounting: the accounting
plan returns None both for an absent key and for a key outside the data
subspace, so deriving the flag from it would have reported that deleting
a metadata key removed nothing. Split, with a test.
Refs #12 feat(node): enforce per-tenant storage quotas and rate limits
Two limits per tenant: stored bytes, and operations per second. Both
default to unlimited, because a node that started enforcing a limit
nobody configured would reject writes for reasons its operator never
chose.
The unit of a quota is the caller's key plus the caller's value. Not the
encoded key: the tenant-name length prefix and the subspace byte are the
node's framing, and billing a tenant for the length of its own name would
be indefensible. Not the engine's on-disk footprint either, because that
moves with compaction and a quota that drifts under the tenant's feet is
not a quota.
Usage is tracked in memory, because a quota check on the write path
cannot afford a scan, and materialised into the tenant's stats subspace
so an operator or a billing job can read it without asking the node.
Startup recomputes it by scanning the real data rather than trusting the
stored total. That costs nothing — the TTL index needs the same scan, and
the two now share one pass — and it makes the counter self-healing: drift
from a crash between the data write and the counter write, or from a bug
in the accounting, is corrected at the next boot instead of compounding.
A write reads the current entry to charge the difference rather than the
whole value. Counting only additions would inflate usage on every
overwrite, and a quota that wanders upwards without the data growing
rejects legitimate writes for a reason the tenant cannot see. The check
runs before the WAL append, so a refusal cannot leave a record replay
would re-apply.
A write that frees bytes is always admitted, even for a tenant already
over its limit. Otherwise the only way out of an exceeded quota would be
an operator raising it, and a tenant could not fix its own problem.
Rate limiting is admission control, so it runs before any storage access
and covers reads as well as writes: a tenant can saturate a node with
scan just as effectively as with put. A batch costs one operation per op
it contains, since charging per RPC would let any client willing to batch
bypass the limit entirely. A request costing more than the whole
per-second allowance is InvalidArgument rather than ResourceExhausted —
no wait admits it, so reporting it as a rate limit would send the client
into a retry loop that cannot succeed.
governor's keyed limiter would allocate a key per request just to look
one up, on the hottest path in the node. The map is ours, keyed by
Arc<str> so it can be probed with a bare &str; the limiters are
governor's.
Admin API on the existing admin server rather than as gRPC RPCs. That
port is already the operator surface and binds to loopback by default,
whereas an administrative RPC on the tenant-facing API would need an
admin RBAC role that does not exist yet.
Closes #26 feat(node): stream key changes over Watch (#52)
* fix(core): reject a zero-length tenant in KeyCodec::decode
`encode` refuses an empty tenant, but `decode` happily read a leading
length byte of zero and handed back `""` as the tenant name. The two
directions disagreed on what a valid key is, so a corrupt or forged key
could decode into a tenant nobody owns — and every caller downstream
would treat that empty string as legitimate.
Found by the `key_roundtrip` fuzz target added in the following commit,
which asserts exactly this symmetry. The regression test is the
deterministic version of what the fuzzer minimised to.
* test(fuzz): add cargo-fuzz targets over every byte-parser
Five coverage-guided targets, one per format the node reads from
somewhere it does not control: keys and values off disk, a WAL segment
after a crash, an operator's TOML.
The targets assert properties rather than only absence of panics. A
decoder that returns Ok with the wrong slice boundaries is a
tenant-isolation bug, and no amount of not-crashing would surface it, so
`key_decode` checks that the decoded parts account for every byte and
`key_roundtrip` checks that decode and encode agree on which keys are
valid. That second assertion is what found the empty-tenant bug fixed in
the preceding commit; with the fix reverted, libFuzzer minimises to it in
under a second.
`value_decode` pins `is_expired` against `decode`, since the fast path
re-parses the TTL header without allocating and two parsers over one
format drift. `node_config` runs TOML through to `validate`, because a
config that deserializes but describes an impossible node has to be
refused with a message rather than an overflow later on.
The seed corpus is hand-written and committed read-only under
`fuzz/seeds/`: valid inputs for every branch of each format plus the
near-misses each parser must reject. A fuzzer derives a well-formed WAL
frame on its own only after rediscovering CRC32, so handing it one is
worth far more than the bytes cost. libFuzzer writes new inputs to the
working corpus only, so CI cannot churn the seeds.
Property tests complement the fuzzing from the other side: generated
*valid* logs, which a mutator reaches slowly and by accident. They pin
that truncating a segment anywhere yields a prefix of the records, that
corrupting a byte can only cost the tail and never resurrect a record,
and that reopening a segment agrees with parsing its bytes directly.
CI runs a 60s pass per target on pull requests and 600s nightly, with
the crashing input uploaded on failure — without it a failed run says
only that something panicked.
Avro schema parsing is in the issue but has no target, because Avro
validation itself is not implemented yet (#27); the target belongs in
that change.
Closes #35
* ci(fuzz): pin the target triple and stop five jobs racing one cache
Every fuzz job failed on the runner while passing locally. cargo-fuzz
derives its target from the host triple, the runner reports a musl host,
and AddressSanitizer cannot be linked against a static libc — the build
died with "sanitizer is incompatible with statically linked libc" before
trying a single input.
The triple is now pinned to x86_64-unknown-linux-gnu and installed
explicitly, with the reason recorded next to it so the next person to see
a musl host does not rediscover it. Noted in fuzz/README.md too, since
the same trap is reachable locally.
Separately, all five matrix jobs shared one cache key, so four of them
failed the reservation and logged an error that reads like a real
failure. One writer, four readers.
* feat(node): stream key changes over Watch
The Watch RPC returned Unimplemented. It now streams puts, deletes and
TTL evictions on a key prefix, off a post-commit hook.
The issue called for a broadcast channel per watched prefix. That is the
wrong key: a prefix is an arbitrary byte string chosen by the client, so
a registry keyed by prefix has to be walked on every write to find the
matching entries and garbage-collected when the last subscriber of each
prefix leaves — a prefix tree maintained on the write path, in exchange
for narrowing a fan-out that is already narrow. Keying by tenant gives a
bounded, meaningful key: a write wakes only the watchers of its own
tenant, which is the isolation boundary the node already enforces, and
prefix filtering moves to the subscriber where it costs one starts_with
on an event that subscriber was going to be woken for anyway.
Tenant isolation is structural rather than checked: the channel is the
tenant, so there is no code path along which an event could cross. The
widest subscription a client can ask for — an empty prefix — is the test
for it.
Publishing cannot fail a write. The hook takes no Result, runs after
storage has accepted the write, and skips the work of building an event
when nobody is watching, so a node with no subscribers pays one hash
lookup per write and the channel map stays empty. A channel is collected
when its last subscriber leaves, on the next publish: a dropped Receiver
has no hook to run, and publish is the next moment anyone looks.
A subscriber that outruns its buffer ends with DataLoss rather than
silently skipping. Skipping is the dangerous failure: the client would
believe it had seen every change and act on a keyspace it never read.
The message names the count and tells it to re-read and resubscribe.
TTL eviction now goes through the service rather than straight to
storage, so it gets a WAL record and is announced. A key that vanishes
on its own is the one change a watcher cannot discover for itself.
The heartbeat is measured as an absolute deadline, not a fresh delay per
loop iteration. The filtered branch loops without yielding, so a relative
timer is reset by every non-matching event: a subscriber on a narrow
prefix in a busy tenant would go permanently silent, precisely when it
most needs to know the stream is alive.
WatchEvent.kind becomes a named enum in the proto. It is wire-compatible
with the uint32 it replaces, and it makes adding a kind a visible change
to the contract rather than a new magic number.
Closes #25 Latest Branches
-1%
-15%
-14%
© 2026 CodSpeed Technology