Skip to content

feat(service): Add CQL high-volume backend for Cassandra and ScyllaDB - #675

Draft
lcian wants to merge 1 commit into
mainfrom
feat/cql-high-volume-backend
Draft

lcian wants to merge 1 commit into
mainfrom
feat/cql-high-volume-backend

Conversation

@lcian

@lcian lcian commented Oct 8, 2026

Copy link
Copy Markdown
Member

Adds CqlBackend, a high-volume backend for Apache Cassandra (5.0+) and ScyllaDB (2025.4+) built on the scylla driver. It can be configured standalone (type: cql) or as the high_volume tier of tiered storage. Nothing changes unless it is configured.

Draft for discussion and sandbox benchmarking — not production-ready. It has only been tested against single-node clusters, and the write path has not been load-tested.

storage:
  type: tiered
  high_volume:
    type: cql
    nodes: [cassandra-1:9042, cassandra-2:9042]
    keyspace: objectstore
    table_name: objectstore
    local_datacenter: us-east1   # optional
  long_term:
    type: gcs
    bucket: my-bucket

The keyspace and table must exist up front; the schema is in devservices/cql-schema.cql and in the CqlConfig docs. Optional settings cover username/password, rustls TLS (incl. mTLS), request timeout, and consistency levels (defaults: local_quorum / local_serial).

Design

Each object is one partition keyed by its storage path, with kind (object / tombstone / upload marker), metadata, payload, redirect, expires_at, and a rev uuid that changes on every write.

CQL IF conditions only support AND, so "absent, inline, or expired tombstone" can't be checked server-side the way Bigtable's CheckAndMutateRow does. Every conditional operation instead reads the row, checks the precondition locally, and writes with IF rev = <observed> (a missing row matches IF rev = null), retrying a few times if it loses a race. Every write, including standalone puts and deletes, is an LWT. Mixing LWT and plain writes on the same row can reorder them under clock skew.

What to look at when benchmarking

  • Write cost is the main open question. Every HV write is a Paxos round (about 4 round trips on Cassandra, about 3 on ScyllaDB) plus a preceding read, and Paxos also persists the payload (up to 1 MiB). Expect noticeably higher write latency and write amplification than Bigtable.
  • Reads are a single LOCAL_QUORUM select; metadata-only reads skip the payload column.
  • Objects near the 1 MiB tiering threshold will trigger ScyllaDB's default large-cell warnings.

Compatibility notes

  • 2038 limit on Cassandra: with its default storage_compatibility_mode, Cassandra 5.0 rejects TTLs that expire after 2038-01-19. Rows with such deadlines, or deadlines more than 20 years out, are written without a TTL. They still expire on read but are only reclaimed when overwritten or deleted. ScyllaDB accepts these TTLs.
  • ScyllaDB tablets: LWT on tablet keyspaces (the default) requires ScyllaDB 2025.4. On older versions, create the keyspace with tablets = {'enabled': false}.
  • Different LWT results: Cassandra and ScyllaDB return different columns from conditional writes, so only [applied] is read.

Known gaps

  • Only tested against single-node clusters. Failure handling for writes with an unknown outcome (re-read and check our revision) is not exercised.
  • TLS has only unit tests, including the custom verifier behind verify_hostname: false.
  • Exhausting CAS retries under heavy contention on a single key returns an internal error.
  • The only metric is cql.failures; driver-level retries aren't counted.

Testing

The Bigtable backend tests are ported (except legacy-format ones), plus CQL-specific tests for TTLs, upload markers, and concurrent CAS. CI runs them against Cassandra in test-all and against ScyllaDB in a new test-cql-scylla job. A Cassandra devservice (127.0.0.1:8088) is added to full mode. apply_range moves from bigtable.rs to common.rs so both HV backends share it.

Add CqlBackend, which implements HighVolumeBackend on top of the scylla driver and works with both Apache Cassandra 5.0+ and ScyllaDB 2025.4+. It is available as a standalone 'cql' storage type and as the high-volume tier of tiered storage.

CQL conditions cannot express the disjunctions that tombstone-aware operations need, so every conditional operation reads the row, checks its precondition locally, and writes with 'IF rev = ?' on the observed revision. All writes are lightweight transactions so they are never reordered against plain writes. Rows carry an authoritative expires_at column plus a TTL for reclamation; deadlines beyond Cassandra's 2038 limit are stored without a TTL.

Also adds a Cassandra devservice, runs the CQL tests against Cassandra and ScyllaDB in CI, and moves apply_range into common.rs so both high-volume backends share it.
@codecov

codecov Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 86.79245% with 7 lines in your changes missing coverage. Please review.
✅ Project coverage is 91.48%. Comparing base (0abc6bf) to head (f0fa574).

Files with missing lines Patch % Lines
objectstore-server/src/config.rs 89.18% 4 Missing ⚠️
objectstore-service/src/backend/mod.rs 0.00% 2 Missing ⚠️
objectstore-test/src/server.rs 0.00% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #675      +/-   ##
==========================================
+ Coverage   91.38%   91.48%   +0.10%     
==========================================
  Files         117      118       +1     
  Lines       24056    25697    +1641     
==========================================
+ Hits        21983    23510    +1527     
- Misses       2073     2187     +114     
Components Coverage Δ
Rust Backend 94.74% <88.46%> (-0.15%) ⬇️
Rust Client 81.70% <ø> (ø)
Python Client 93.81% <ø> (ø)

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@lcian

lcian commented Oct 8, 2026

Copy link
Copy Markdown
Member Author

bugbot run
@sentry review

Comment on lines +119 to 126
/// CQL backend for [Apache Cassandra] or [ScyllaDB].
///
/// [Apache Cassandra]: https://cassandra.apache.org/
/// [ScyllaDB]: https://www.scylladb.com/
Cql(cql::CqlConfig),
}

/// Constructs a type-erased [`common::HighVolumeBackend`] from the given config.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: The compare_and_update method for the CQL backend lacks a retry loop for its CAS operation, unlike other conditional write methods in the same file.
Severity: HIGH

Suggested Fix

Wrap the logic inside compare_and_update in the CqlBackend with a retry loop, similar to the pattern used in other conditional write methods like put_non_tombstone. On a failed CAS attempt (when write_if returns false), the loop should retry the read-modify-write cycle instead of immediately returning SetExpiryResponse::Rejected.

Prompt for AI Agent
Review the code at the location below. A potential bug has been identified by an AI
agent. Verify if this is a real issue. If it is, propose a fix; if not, explain why it's
not valid.

Location: objectstore-service/src/backend/mod.rs#L119-L126

Potential issue: The `compare_and_update` method in `CqlBackend` performs a client-side
read-then-write check for its Compare-And-Set (CAS) operation but lacks a retry loop to
handle race conditions. If a concurrent write occurs between the read and the
conditional write, `write_if()` will fail, and the method will incorrectly return
`SetExpiryResponse::Rejected`. This signals a permanent failure instead of a transient
one that could be retried. As a result, under concurrent load, updates to object
expiration times can be silently lost, leading to objects having stale expiration data.

Did we get this right? 👍 / 👎 to inform future reviews.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want reviews to match your repository better? Bugbot Learning can learn team-specific rules from PR activity. A team admin can enable Learning in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit f0fa574. Configure here.

));
}
Ok(())
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Upload marker create is not idempotent

Medium Severity

create_upload_marker writes only when the row is absent (IF rev = null). A retry after a successful create, or a leftover expired marker at the same upload key, loses the LWT and returns an internal error. The Bigtable backend overwrites unconditionally, so the same retry succeeds there.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit f0fa574. Configure here.

)
.await?;
applied(result)
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Marker delete misses timeout recovery

Medium Severity

delete_upload_marker returns the driver error when the LWT times out and never re-reads the row. write_if and delete_if treat a matching post-error state as success. A delete that applied before the timeout then looks like a missing marker on retry, so upload completion can fail after the marker was already consumed.

Additional Locations (1)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit f0fa574. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant