\This is the complete Casita documentation, as Markdown.
# Introducing Casita: A content-addressed store for source code and build artifacts
> Source code and build artifacts sharing verified storage, synchronization, and garbage collection.
We’re rewriting [Nix](https://nix.dev/) in Rust, and it needs to be [split into layers](/concepts/responsibilities/) and modernized.
Its storage, build machinery, and higher-level tools can each be useful on their own.
You should be able to use one layer without adopting the whole stack.
That matters even more in agentic development. Agents produce more source code, more versions of it, and more build artifacts as they explore and test changes.
Keeping every copy quickly becomes expensive, but throwing everything away means rebuilding or regenerating work you might need again.
Developers are sharing screenshots of disks filled with `target/` directories in a matter of hours.
It feels familiar to how Nix users who have had to garbage collect their `/nix/store`: which artifacts are still useful, which ones can go, and how to reclaim space without losing what another project still needs.
Casita is our first standalone layer, a [content-addressed object store](/overview/) for source code and build artifacts, with [shared storage](/concepts/deduplication/), [verification](/concepts/verification/), [synchronization](/concepts/sync/), and [garbage collection](/concepts/garbage-collection/).
It is still pre-release and available as a Rust library and CLI. We’re targeting Linux, macOS, and Windows.
## The storage problem in ~~Cargo~~ package managers
Rust workspaces and throwaway checkouts can accumulate `target/` directories. Deleting them discards artifacts a later build might reuse; keeping them duplicates bytes across similar projects. [Cargo](https://doc.rust-lang.org/cargo/guide/cargo-home.html) and [uv](https://docs.astral.sh/uv/concepts/cache/) cache downloaded dependencies, but their caches do not manage source versions and generated outputs across projects.
We added an [`ArtifactStorage` interface](https://github.com/cachix/cargo/commit/a393d0143a6a5ee600e2c6f87d0eb537de8954d2) so Cargo can prepare and persist registry archives, Git dependencies, and workspace build outputs through different backends. The filesystem backend preserves Cargo’s usual behavior.
The [Casita backend](https://github.com/cachix/cargo/commit/7953f9427a3dba5eae5f3eb687d7a8ae9b78fa3b) imports and restores that content through [local IPC](/integrations/ipc/), while Cargo still decides what to download and build. This remains experimental: import and restore performance has not yet matched the filesystem backend. The [development branch](https://github.com/cachix/cargo/tree/artifacts%2Bcasita) and [Cargo guide](/integrations/cargo/) have the details.
### Cargo using Casita
WriteRun
Pause
Cargo writes all three outputs directly into Casita.
ABCACDEABC
### Cargo
$ cargo build --bin uv
**libuv.rlib**
**ACD→
Waiting
**uv**
**ABC→
Waiting
**uv.d**
**E→
Waiting
Writing outputs
**
### Casita
**cargo/builds/uv**Named root
**Output directory**Directory object · file entries
**uv**File object
**libuv.rlib**File object
**uv.d**File object
**Blob**File contents
**Blob**File contents
**Blob**File contents
**A**chunk
**B**chunk
**C**chunk
**D**chunk
**E**chunk
Storage links, not compilation dependencies. Chunk boundaries and sharing are illustrative; directory payload details are omitted.
### Run uv
**uv**
ABC
Awaiting metadata
$ casita run cargo/builds/uv -- --version`uv 0.12.7 (x86_64-unknown-linux-gnu)`
Output tree prepared. Executable discovered and run.Replay
Proposed Cargo integration · uv 0.12.7 · executable reconstruction shown; chunks are illustrative. [Run applications →](/guides/run/) · [Verified artifact ↗](/examples/uv-artifacts.json)
## How Casita works
Think of Casita’s storage model as a generalized Git object database. Git has blobs, trees, and commits. Casita stores immutable byte blobs and immutable object records: each format defines an object’s identity and its links to other objects. A directory links to its files and subdirectories; a Git commit links to a tree and its parents. Casita follows those links to find a complete saved version.
Every blob has a [BLAKE3](https://github.com/BLAKE3-team/BLAKE3/blob/master/README.md) hash of its complete bytes. BLAKE3 computes that hash as the root of a Merkle tree. Casita can keep an optional [Bao *outboard*](https://github.com/oconnor663/bao#outboard-mode) with the tree’s intermediate hashes, so a reader can verify one range against the blob’s hash without reading the whole blob. A full sequential read checks the complete hash. The storage backend may chunk and compress the bytes without changing their address.
An object record gives those bytes meaning and links. For `src/main.rs`, a file record points to the blob containing its source code. The `src/` directory has its own record, which links to that file record and points to a blob encoding the directory entries. These records form a graph above the blobs.

For this [filesystem tree](/concepts/directory-storage/), the hashes connect roughly like this:
```text
main.rs ID = BLAKE3(file bytes)
src/ ID = BLAKE3(canonical entries containing main.rs ID)
```
Changing `main.rs` creates a new file ID, which changes the `src/` directory ID and the IDs of its parent directories. Unchanged files keep their IDs and stored bytes. The old blobs and records are never rewritten.
The blob’s BLAKE3 hash and a source format’s hash serve different purposes. A [Git commit](/guides/git/) keeps its native Git ID while Casita addresses its body bytes by their BLAKE3 hash. A [Nix archive (NAR)](https://nix.dev/manual/nix/2.35/command-ref/nix-store/dump) has a SHA-256 hash of its canonical serialization. Casita measures that NAR hash, then stores the archive’s files as a graph of BLAKE3-addressed blobs. The verified object records let Casita traverse the graph without decoding every payload.
Read [Blob Storage](/concepts/blob-storage/) and the [repository model](/concepts/repository/) for the detailed contracts. The [NAR IPC guide](/integrations/ipc/#nar-and-filesystem-nar) shows how to import and restore an archive.
## Repository workflows
### Roots and retention
An application gives a saved graph a name, called a [*root*](/concepts/roots-and-retention/). The root points to one exact object and keeps everything reachable from it. The objects and blobs stay immutable; the application can move or remove the root as its needs change.
Applications can save separate versions explicitly. For example, import a project directory after two revisions under `projects/app/v1` and `projects/app/v2`. The imports leave the original working directory in place. Both names keep their versions available, and identical files share one blob. An application could also point `projects/app/current` at the same object as `v2`, then move that name to a later version without changing either saved graph.
Casita follows links from a root to find its complete graph for synchronization and retention. Removing a name makes objects needed only by that name eligible for garbage collection once active work releases them; files and chunks shared with another root remain.
Roots are permanent by default, which suits saved Git histories and releases. Rebuildable data, such as a Cargo `target/` directory, can instead use an evictable root:
```bash
casita import ./target --root cargo/my-app/target --retention evictable
casita root ls cargo/my-app/target --long
```
The CLI can also change retention later with [`root retention`](/reference/cli/#root-retention). The [Rust API](/reference/rust-api/) provides `set_root_with_retention` and `touch_root` for applications that manage their own cache roots. Marking a root evictable does not remove it immediately.
### Filesystem, tar, NAR, and Git imports
[Importers](/concepts/imports/) turn inputs into graphs that Casita can verify and retain. The [filesystem importer](/guides/filesystem/) walks a directory; the [tar importer](/guides/tar/) reads an archive without extracting it first; and the [NAR importer](/integrations/ipc/#nar-and-filesystem-nar) measures a Nix archive’s canonical SHA-256 while storing its files as a graph. The [Git importer](/guides/git/) keeps native Git object IDs and a view of selected branches and tags.
Each importer uses the same repository publication and retention rules. Once content is saved under a root, Casita can synchronize it and collect it when no name or active work needs it.
### Casitar archives
[Casitar](/guides/casitar/) writes a complete saved graph to a portable `.casitar` archive. You can pass it through a file, pipe, or release artifact when the source repository is unavailable to the receiver. On import, Casita checks every payload and object record, verifies that the archive contains the whole graph, and then publishes the destination roots together. The archive carries object identities and bytes, independent of the source repository’s packing and database layout.
### Git object database adapter
The experimental [Gix object database adapter](https://github.com/cachix/casita/blob/main/crates/casita/examples/gix_casita_odb.rs) reads and writes native Git objects through Casita. Git’s SHA-1 or SHA-256 IDs stay intact while Casita verifies and stores the object bodies and their links. The adapter covers object storage; a complete Git repository also needs its refs, index, and working tree integrated. The example writes a blob, tree, and commit, then reopens the store and reads the commit back.
### CasitaFS mounts
The companion [CasitaFS](https://github.com/cachix/casita/blob/main/crates/casita-fs/README.md) crate presents a saved filesystem tree as a read-only mount. Existing tools can browse its files without first checking out the whole tree. It uses FUSE on Linux and a native FSKit extension on macOS 26 or later. The macOS extension requires explicit setup for each user.
### Repository sync
Suppose a second repository already has `projects/app/v1`. Syncing `projects/app/v2` reuses what is there and sends missing content. Compatible stores can reuse chunks within changed files too. The destination checks incoming objects against their format’s rules and verifies that the whole saved version arrived before it updates the name. An interrupted transfer can be retried without exposing a partial version through that name.
Local repository sync is implemented in the CLI and Rust library. An optional SSH source uses OpenSSH for the connection while the receiving Casita repository still verifies what arrives. Sync adds content; removing destination names and collecting unused data are separate actions. See [Synchronization](/guides/sync/) for local, SSH, and selective workflows.
### Garbage collection
If both `projects/app/v1` and `projects/app/v2` are named, Casita keeps both. Remove the first name and run collection: bytes needed only by `v1` can go, while files and chunks shared with `v2` stay. Active readers and writers also hold on to the data they are using during collection.
The local CLI can preview a pass with `gc --dry-run` before running `gc`.
Automatic collection helps keep a busy disk from filling up. In the standard local repository, starting a mutation when the filesystem is at least 80% full triggers a nonblocking collection attempt. Casita first reclaims data no root or active operation needs. If disk use remains at or above 75%, it releases the least recently used roots explicitly marked evictable, vacuuming after each release until use falls below 75% or no eligible roots remain. Permanent roots stay, and manual `gc` preserves every named root, including evictable ones.
Collection applies to the Casita repository. See [Garbage Collection](/concepts/garbage-collection/) for the retention and recovery rules.

### Shared S3 repository
Two runners can use the same S3 bucket and prefix to publish objects and read named versions. Their payload bytes are immutable, but names and object records change as new versions arrive. Casita records those changes in Chroma’s **wal3**, a write-ahead log stored in S3. Conditional updates to its manifest give competing writers an agreed order for their changes.
Readers and writers register durable holds so collection preserves data still in use. The S3 profile is experimental and requires the `s3` feature; an abandoned hold needs explicit recovery. The [shared S3 guide](/guides/s3-multi-owner/) shows how independent owners can use one repository, and [S3 maintenance](/guides/s3-maintenance/) covers collection and recovery.
## What’s left for the 0.1 release
Casita runs from a source checkout today, but we have not tagged `v0.1.0`. Before the first release, we need to improve performance where benchmarks show bottlenecks, broaden benchmark coverage, and use Casita in more real projects and end-to-end workflows. That practical use should expose problems we can fix before calling 0.1 ready.
The integrations have more work ahead. Cargo’s Casita backend needs faster import and restore before it can match the filesystem backend. The Git object store adapter still needs integration with the rest of a Git repository.
## Try the current pre-release
From a Casita source checkout, install the CLI, then import a directory into the default repository, list its roots, and preview collection:
```bash
cargo install --path crates/casita
casita import ./src --root examples/source
casita root ls
casita gc --dry-run
```
When roots and object records are in one Casita repository and the matching blobs are in another, choose the blob source during sync:
```bash
casita sync --from ./metadata-store \
--from-blobs ./blob-store \
--to ./mirror \
--root examples/source
```
The [sync guide](/guides/sync/#read-payloads-from-another-repository) explains the requirements for a separate blob source.
The [Quick Start](/getting-started/) walks through checkout and inspection, and the [Library guide](/library/) shows the supported Rust API.
Our [filesystem benchmark](/benchmarks/) compares imports with Git add and commit. In that run, Casita imported the large-file corpus faster, while Git was faster for small-file imports and unchanged re-imports. The [full benchmark reference](/reference/benchmarks/) includes methodology and separate native Git measurements.
# CLI
> Common commands for storing, copying, and checking objects.
Install the command from a [source checkout](../getting-started/):
```console
$ cargo install --path crates/casita
```
The examples use `--repository ./cache` to keep data in a directory you choose. Without that option, Casita uses its per-user data directory. `sync` uses `--from` and `--to` instead.
## Store and read files
```console
$ casita --repository ./cache import ./project --root projects/demo
$ casita --repository ./cache root ls
$ casita --repository ./cache tree list casita.directory.v1:...
$ casita --repository ./cache checkout casita.directory.v1:... ./restored
```
Replace the example key with the one printed by `import`. `tree list` shows a directory’s direct entries; `checkout` restores it to an absent or empty path. See [filesystem import](../guides/filesystem/) for cache and checkout behavior.
## Keep or release a graph
```console
$ casita --repository ./cache root set releases/current casita.directory.v1:...
$ casita --repository ./cache root rm releases/current
$ casita --repository ./cache gc --dry-run
$ casita --repository ./cache gc
```
A root name retains the complete graph it points to. Removing a name makes unshared data eligible for collection; it does not immediately delete bytes. Checkout creates its own `auto/checkout/...` root unless you pass `--no-root`.
## Copy a graph
```console
$ casita sync --from ./cache --to ./mirror --root projects/demo
```
The destination verifies received objects before moving the selected root. For SSH and S3 endpoints or copying a single path, see the [sync guide](../guides/sync/).
## Check integrity
```console
$ casita --repository ./cache fsck --audit-only
$ casita --repository ./cache fsck --dry-run
$ casita --repository ./cache fsck --source ./verified-replica
```
The first command only audits. The second previews safe physical repairs; the third may repair physical data from an independently verified local replica. See [operations](../guides/operations/) before restoring or backing up a repository.
## Other workflows
* [Run a stored application](../guides/run/)
* [Import tar](../guides/tar/), [Git](../guides/git/), or [Casitar](../guides/casitar/)
* [CLI reference](../reference/cli/) for every option and its exact behavior
# Concepts
> How objects, roots, verification, storage, and collection fit together.
Casita stores objects and their links as immutable graphs. A format decides what an object means and how to verify its identity. A named root points to one object and keeps its entire reachable graph available.
## Follow an object through the repository
1. An [import](./imports/) reads data and creates verified object records.
2. A [root](./roots-and-retention/) names the graph to keep.
3. Casita stores payload bytes, possibly using chunks and [deduplication](./deduplication/). Those storage choices do not change object identity.
4. [Sync](./sync/) copies selected objects or roots to another repository. The destination [verifies](./verification/) what it receives.
5. [Garbage collection](./garbage-collection/) removes objects that no root or active operation still needs.
For example, a Git commit keeps its Git ID, an IPLD block keeps its CID, and a filesystem directory gets a canonical directory key. All three use the same root, sync, and collection machinery.
## Two kinds of identity
| Term | What it identifies |
| ----------- | ----------------------------------------------------------------------------- |
| `ObjectKey` | One object in a format’s namespace, such as a Git object or Casita directory. |
| `BlobId` | The complete plaintext bytes stored for an object, using a BLAKE3 digest. |
An `ObjectRecord` connects the two: it records the object key, payload ID and length, and exact forward links. Several objects can share one payload. Read [repository](./repository/) for the full state model and [blob storage](./blob-storage/) for physical layout.
Roots and repository revisions belong to one repository. They do not provide distributed ordering for names copied between repositories.
## Go deeper
* [Directory storage](./directory-storage/) explains canonical filesystem trees.
* [Shared payload services](./shared-payload-services/) covers several repositories sharing bytes without sharing deletion authority.
* [Responsibilities](./responsibilities/) shows which decisions applications make around Casita.
# Blob Storage
> How Casita identifies, stores, and reads immutable payload bytes.
A **blob** is a byte sequence identified by the BLAKE3 digest of its complete plaintext, its `BlobId`. Equal bytes have the same ID even if different importers wrote them or stores chunked and compressed them differently.
`BlobStore` supplies presence checks, readers, streaming writes, and optional chunk information. The standard native store uses FastCDC chunks compressed with zstd. A manifest binds those chunks into one blob; a single-chunk blob can omit the manifest because its chunk ID already equals the blob ID.
## Read guarantees
The standard chunked store checks a chunk’s digest before returning its bytes. An unseeked sequential read also checks the complete `BlobId` at EOF. A caller that stops early has not completed that whole-blob check. Use `Repository::open_verified` when bytes must be authenticated before they are returned. It reads a Bao proof and fails if that proof is unavailable or invalid. Bao also supports checking a selected range without reading the rest of the blob. See [Verification](../verification/) for the distinction.
Roots and collection belong to `Repository`, not `BlobStore`. The payload backend alone cannot decide when a blob is safe to delete.
## Put a near tier in front of a far tier
With the `experimental` feature, `CombinedBlobStore::new(near, far)` reads from the writable near store first and falls back to the far store when a blob is absent. Writes go only to near, and reads from far do not fill near.
```rust
use casita::experimental::{CombinedBlobStore, MemoryBlobStore, MemoryMetadataStore, Repository};
fn tiered_repository() -> Result<
Repository, MemoryMetadataStore>,
Box,
> {
let near = MemoryBlobStore::new();
let far = MemoryBlobStore::new();
let payloads = CombinedBlobStore::new(near, far);
Ok(Repository::new(payloads, MemoryMetadataStore::new()?))
}
```
The adapter has no `BlobGc` implementation because the repository cannot know who else still needs data in a shared far store. It also has no combined `BlobSync`, so transfer uses whole payload streaming. A repository using a shared payload service can prune its own records with `collect_logical`, while the service controls physical deletion. See [Shared Payload Services](../shared-payload-services/) for that contract.
### Repair a near tier from a replica
`RepairingBlobStore::new(near, far)` accepts two `ChunkedBlobStore` values. It checks the complete near blob before returning an ordinary reader. If near has typed corruption or a referenced chunk is missing, it verifies the complete far blob, repairs near, and checks the replacement. Permission and backend errors do not trigger repair. A failed repair retains both the near error and the repair error in `BlobRepairError`.
`verified_read` can also rebuild missing or corrupt Bao data from verified near bytes. The adapter collects only near data; the far replica needs its own retention policy. Its extra full read makes it suitable when verified repair matters more than read latency.
Read [Deduplication](../deduplication/) for chunk reuse and the [Experimental Rust API](../../reference/experimental-rust-api/#payload-storage) for backend methods.
# Deduplication
> How equal blobs and shared chunks reduce storage and transfer work.
Casita reuses content at two levels. Equal complete byte sequences share one `BlobId`. For similar blobs, the standard store uses FastCDC to find chunk boundaries in the plaintext. Its default target is 256 KiB, with a 128 KiB minimum and 512 KiB maximum.
A small edit may shift nearby boundaries, but FastCDC usually finds the same boundaries later in the file. The new blob then stores changed chunks and reuses unchanged ones. Compression happens after chunking, so it does not change the blob’s identity or those boundaries.
Chunk IDs describe physical storage, not logical objects. A store can change its chunk size without changing any `BlobId` or `ObjectKey`. Repositories with different chunk settings can still exchange the same blobs.
During sync, compatible chunk stores can negotiate which chunks the destination already holds and send only missing chunks. Other transfers stream the complete plaintext payload. Collection keeps a shared chunk while any live blob still references it.
See [Blob Storage](../blob-storage/) for reads and physical layout, and [Sync](../sync/) for transfer choices.
# Directory Storage
> How canonical directories link files and subdirectories into a graph.
A directory maps validated names to three kinds of entry:
| Entry | Stored information | Retaining link? |
| ------------ | -------------------------------------- | --------------------------- |
| File | Blob ID, byte size, executable bit | Yes, to the blob |
| Subdirectory | Directory ID and recursive entry count | Yes, to the child directory |
| Symlink | Validated target stored inline | No |
The encoding fixes entry order and field representation. Equal directory contents therefore produce the same `DirectoryId`, which is the BLAKE3 digest of the canonical payload. Timestamps, ownership, ACLs, and extended attributes are not part of that identity.
A directory is an ordinary `casita.directory.v1` object. Its format verifier checks the payload, key, and exact child links. Closure verification also checks child relationships, including the declared file sizes and recursive counts. A missing or inconsistent child prevents the directory graph from being rooted as complete.
Directories and their children form a Merkle graph. A directory’s **closure** is the directory plus every reachable subdirectory and file blob. Roots retain that closure, sync copies it, and collection keeps it while the root or another hold is live.
See [Filesystem Imports](../imports/#filesystem-imports) for capture behavior, [Blob Storage](../blob-storage/) for file bytes, and [Sync](../sync/) for transfers.
# Garbage Collection
> What collection keeps, how to run it, and how it recovers.
Casita collects objects that no named root or active operation needs. You can run collection explicitly. The standard local profile also attempts a nonblocking pass before a mutation when disk usage reaches 80%. If that pass leaves disk use at or above 75%, it releases local roots explicitly marked evictable, oldest used first, and vacuums after each release. Permanent roots are never selected. A completed pressure pass has a 60-second cooldown shared by local repository handles.
## What stays live
Each pass marks one logical-state snapshot. It keeps every named root’s forward closure and the data protected by active pins, including reads, transfers, and unpublished writes. It also keeps the payloads and chunks needed by those objects. Shared bytes survive while any retained graph uses them. Unrooted data is allowed and becomes collectible after its pins end.
Collection runs alongside readers and writers; only collectors serialize with one another. A large traversal spills temporary state to `\/spill` rather than requiring the entire graph in memory. Spill files are deleted after the pass, and abandoned files are cleaned up on the next open. `SpillLimits` bounds temporary memory and disk use.
## Preview and collect
Use the CLI to inspect one collection plan and then run a fresh pass:
```console
$ casita --repository ./cache gc --dry-run
$ casita --repository ./cache gc
```
`--dry-run` computes a plan without changing state. A later `gc` computes a fresh plan, so counts can change if roots or objects changed meanwhile. Manual `gc` and `vacuum` do not evict named roots, including evictable ones.
Library callers have three choices:
```rust
let preview = repository.preview_collection().await?;
let outcome = repository.collect().await?; // waits for another collector
let outcome = repository.try_collect().await?; // returns Busy instead of waiting
```
`collect` waits for another collector. `try_collect` returns `Busy` if ownership or a racing pin prevents the pass; a scheduler can retry later. The local pressure check uses nonblocking `try_vacuum` and lets the mutation continue when collection is busy.
## Shared payload services
Several runners opening the same S3 repository share one root set and one collector. Its [S3 maintenance guide](../../guides/s3-maintenance/) covers hold inspection and recovery.
Separate repositories sharing one physical payload service need a different rule: one repository’s roots cannot decide when shared bytes are deleted. Each repository can prune only its own logical records:
```rust
let preview = repository.preview_logical_collection().await?;
let outcome = repository.collect_logical().await?;
```
Logical collection does not delete payloads or chunks. The shared service must track ownership across all repositories before reclaiming physical bytes. See [Shared Payload Services](../shared-payload-services/) for that contract. `try_collect_logical` gives schedulers nonblocking admission.
## Full-filesystem recovery
Normally, Casita commits the logical prune before deleting physical bytes. If the standard local profile runs out of space during that commit, it can delete only payloads already proven stale by the mark, then retry the prune with the space reclaimed. The pin ledger reserves bookkeeping capacity for these ownership changes.
This fallback works because local state and payloads share one filesystem. A custom repository cannot assume that deleting payloads gives its metadata store space; it returns `StorageFull` instead. Recovery still needs enough reclaimable garbage to fund the logical commit. If none exists, collection returns `StorageFull` without deleting rooted data.
## Commit and sweep order
After taking collector ownership, Casita marks roots and pins, verifies marked records and payloads, atomically prunes unreachable logical records, then deletes unreferenced payload manifests and chunks. Missing rooted data aborts before the prune. A physical deletion failure may leave space to reclaim on a later pass, but does not invalidate rooted data.
The local full-disk fallback above is the one exception to logical-first ordering. If a process stops after its emergency physical sweep, rooted graphs remain valid and another `gc` pass finishes cleanup.
After a local collector crashes, reopen the repository and run `gc` to recover its interrupted claims and prune fence, then audit the recovered state:
```console
$ casita --repository ./cache gc
$ casita --repository ./cache fsck
```
`fsck` may return `Busy` while a collector prevents safe admission. Retry afterward. Remote recovery uses the exact-token procedure in [S3 maintenance](../../guides/s3-maintenance/). `fsck` reports reachable corruption separately from collectible residue; it does not rebuild records from physical bytes.
## Scheduling and storage ownership
Run collection after imports that leave staging data or at a deployment specific disk threshold. `DiskPressurePolicy` controls the local threshold (default 80% used) and calls `try_vacuum`; an external scheduler decides when to retry `Busy`. Zero reported free bytes still triggers an attempt.
When independent repositories share a physical store, aggregate reachability or explicit ownership across all of them. `CombinedBlobStore` alone cannot decide global liveness.
# Imports
> How Casita turns external data into verified objects and a named root.
An import reads an external source, verifies the resulting objects, and publishes them in a repository. A named root keeps the finished graph available. Importers share the repository’s object records, roots, and collection rules; each format does not need its own storage system.
## Publication and failure
Importers stage payloads in a mutation session, which protects those bytes while the import runs. Each object is checked against its format before its record is published. Large imports can publish intermediate batches without a root, then move the named root once the complete graph is ready. If an import fails, those intermediate objects may remain until collection, but the destination root does not point to a partial graph.
The standard local repository checks disk pressure when a mutation starts and may run collection first. See [Garbage Collection](../garbage-collection/) for the policy and explicit cleanup commands.
## Filesystem imports
`FilesystemImport` reads regular files, builds directories from their children, and publishes the resulting directory under a root. For example:
```console
$ casita --repository ./cache import ./project --root projects/demo
```
The walk opens the source directory and resolves descendants relative to that open handle. It refuses a linked root and does not follow symlinks or other reparse points encountered during the walk. A symlink inside the tree is stored as a link with its target, without reading the target.
On supported local platforms, a repeat import can skip reading a regular file when its device, inode, size, modification time, and change time match the last import. This is a speed optimization based on file metadata, not a fresh content check. A remembered result is used only while its object still exists in committed repository state. Platforms without a usable file identity reread every file.
To read and hash every regular file regardless of its metadata, use `--filesystem-rehash` or `FilesystemImport::new(...).reread(true)` in Rust. Read [Capture and Restore a Filesystem Tree](../../guides/filesystem/) for the full workflow.
## Other inputs
| Input | What the importer publishes | Guide |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------- |
| Git repository | Native Git objects and an immutable ref view under `git/\`. | [Git](../../guides/git/) |
| Decompressed tar stream | A canonical filesystem tree, without extracting the archive first. | [Tar](../../guides/tar/) |
| OCI registry image | A single-platform OCI image layout with a bounded config and streamed, digest-verified layers, plus an optional merged filesystem. Requires the `oci` feature. | [OCI](../../guides/oci/) |
| Casitar archive | Verified objects and destination-owned root mappings after the archive’s declared closure checks pass. | [Casitar](../../guides/casitar/) |
IPLD and custom formats also publish through ordinary mutation sessions. See [Add a New Importer](../../guides/adding-an-importer/) for the Rust contract.
# Repository
> How payloads, object records, roots, and verification fit together.
A repository combines payload bytes with revisioned logical state. The supported `casita::Repository` handle provides local and in-memory versions. The generic `casita::experimental::Repository\` lets Rust applications choose backends and a deployment profile.
| Part | Responsibility |
| ------------------- | -------------------------------------------------------------------------------- |
| `BlobStore` | Stores plaintext payloads by `BlobId`; it may chunk or compress them internally. |
| `MetadataStore` | Stores immutable object records, named roots, and the repository revision. |
| `FormatRegistry` | Checks each namespace’s native identity, canonical payload, and direct links. |
| `RepositoryProfile` | Selects coordination, spill placement, and optional maintenance policy. |
`Repository::local(path).await` uses `path/blobs` for payloads and `path/casita.sqlite` for logical state, with lock files for coordination across processes. `Repository::memory()` is temporary. Generic compositions start without local coordination or automatic maintenance; their payload and metadata backends must still preserve the same publication and retention guarantees. See the [Experimental Rust API](../../reference/experimental-rust-api/#deployment-profiles) for custom profiles.
## Logical and physical state
An `ObjectKey` identifies one object in its format’s namespace. Its `ObjectRecord` binds that key to a `BlobId`, plaintext length, and ordered direct links. Several objects may share one payload. A named root points to an object and retains its complete forward closure.
Chunk boundaries, compression, and Bao proof data are physical storage details. Changing them does not change an object key, record, or root.
## Publication lifecycle
1. A mutation session protects payloads while they are staged.
2. The registered format verifies each object before its record publishes.
3. Verified records can publish in bounded batches. A named root moves only when its entire target closure is complete and valid.
4. Retained reads and transfers hold one stable snapshot while collection may reclaim unrelated data.
Failed imports and transfers may leave unrooted objects that collection can later remove. See [Imports](../imports/), [Roots and Retention](../roots-and-retention/), and [Garbage Collection](../garbage-collection/) for those lifetimes.
## The locality boundary
A `RepositoryRevision` compares state within one repository. It is not a clock or a value to order across replicas. A `RepositoryGeneration` orders the states one repository’s readers observe, and is likewise meaningless across repositories. Roots are also local mappings. Sync can copy a source root’s selected value to a destination, but it does not merge concurrent name changes.
`MutationSession::publish_if_roots_match` lets a service publish only while watched roots still have expected values. The service remains responsible for its own writer policy and any ordering it needs across repositories.
# Division of Responsibility
> Which guarantees belong to Casita, object formats, physical backends, higher-level tools, and deployments.
Casita is a repository backbone, not the complete product above it. Its generic contract is strongest when each layer owns only the guarantees it can actually enforce.
## Responsibility map
| Concern | Casita owns | Another layer owns |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Logical identity | A namespace-qualified `ObjectKey`, canonical `ObjectRecord`, and immutable key-to-record binding | The format defines its native identifier and canonical payload |
| Relationships | Stored, verifier-produced forward links used by traversal and retention | The format decides which payload references are true retaining links |
| Publication | Payload durability, namespace verification, bounded staging, and atomic logical commit | The calling workflow decides what should be published and under which destination-owned names |
| Retention | Named roots, complete closure checks, and reachability-based collection | The application decides why a graph matters and when to repoint or remove its root |
| Transfer | One stable source revision, additive batches, receiver verification, and final root publication | Discovery, scheduling, authorization, topology, and destination policy |
| Physical storage | Complete plaintext `BlobId` contract and safe collection ordering | A `BlobStore` chooses chunking, compression, tiers, request behavior, and durability guarantees |
| Logical state | Atomic revisioned records and roots through `MetadataStore` | A deployment decides whether a backend and its backup procedures meet operational requirements |
| Git | Native immutable objects, ref views, checkout, and read-only fetch | Authoring, branches as mutable collaboration state, review, merge, push, and hosting administration |
| IPLD | Registered codec verification, CID identity, and exact links | Selectors, content routing, discovery, and network policy |
| Security | Bounds, receiver verification, safe filesystem containment, typed failure categories, and consumption of already-resolved process credentials | SecretSpec resolves deployment secrets; the deployment owns authentication, multi-tenancy, access control, confidentiality, key management, and audit policy |
For a shared physical payload service, Casita’s logical collection removes only the records in one repository. The deployment owns the aggregate cross-tenant payload ledger, write leases, physical reclamation, and every authorization check. See [Shared Payload Services](../shared-payload-services/).
## Trust boundaries
* **Remote senders are untrusted discovery sources.** A destination reruns its own namespace verifier and publishes a requested root only after its complete closure succeeds.
* **The logical state backend is trusted infrastructure.** Casita detects many inconsistencies with `fsck`, but applications should not mutate database rows outside the `MetadataStore` contract.
* **OpenSSH owns SSH-channel security.** Casita does not weaken host-key verification or replace credential and proxy configuration.
* **Stored content is not automatically confidential.** The standard local profile does not promise encryption at rest or access-pattern hiding.
* **SecretSpec is outside the repository boundary.** The optional S3 profile consumes standard AWS environment variables that `secretspec run` may populate. Casita does not invoke provider vaults or persist resolved values, and does not provide client-side encryption for that profile.
* **A valid claim is not accepted policy.** Evidence formats preserve exact inputs and outputs; the application decides which issuer, time window, or decision is trusted.
## Why the boundary matters
Moving product policy into the generic repository would make object meaning depend on deployment state and would force unrelated formats through one semantic model. Moving repository correctness into each application would duplicate publication, transfer, collection, and recovery logic. The split keeps immutable meaning format-owned and reusable lifecycle machinery generic.
Continue with [Repository](../repository/) for the concrete composition or [Verification](../verification/) for the checks performed at each boundary.
# Roots & Retention
> What keeps an object graph available and how roots change safely.
A `RootName` is a validated hierarchical name pointing to one `ObjectKey` in one repository. Setting or replacing a root commits a new repository revision. The root retains its target and every object reachable through verified forward links.
Removing a root makes that graph eligible for collection if no other root or active hold still needs it. Shared blobs and chunks remain while another live graph references them. Roots are permanent by default. On a local repository, an application can mark a root evictable for disk-pressure collection, which releases least recently used evictable roots when space is low. This is useful for rebuildable artifacts; Git histories and releases can remain permanent. The policy is repository metadata and does not change the root record’s frozen encoding.
## Active operations also retain data
Mutation sessions protect staged payloads until publication finishes or the session ends. A retention hold keeps the snapshot it read usable while other writers and collection proceed. Collection preserves these pinned lifetimes while removing unrelated garbage.
A snapshot hold covers records created through its metadata generation. It does not retain unrelated records published later. Creation generations are repository bookkeeping, not part of an object’s content identity. See [Garbage Collection](../garbage-collection/) for the full collection sequence.
## Conditional publication
`MutationSession::publish_if_roots_match` can publish staged objects and root changes only while every watched name still has its expected target, including expected absence. An unrelated state change is retried; a watched name change returns `RootMismatch` without publishing those logical changes.
This protects a service from overwriting a root that changed since it was read. The service still needs to control who may write its names. The check compares the current target, not the name’s full history: after removal, a delayed create may succeed again. To reject that replay, advance an owner fence root in the same commit. See [Share an S3 Repository Across Owners](../../guides/s3-multi-owner/) for an example.
Roots and revisions have meaning within one repository. `sync --root` installs the selected source value at the destination after verification; it does not order concurrent changes across repositories. Applications needing a merge policy must provide one.
# Shared Payload Services
> Keep physical blobs safe when several repositories share one store.
Ordinary collection assumes that one repository owns its payload store. If several repositories share the same blobs, one repository’s roots cannot say when a blob is safe to delete. Another repository may still have a committed record for it.
```text
repository A ── records ─┐
repository B ── records ─┼─ BlobId ── shared payload service
repository C ── records ─┘
```
Casita’s logical collection removes unreachable records from one repository and leaves physical deletion to the shared service.
## Choose the right collection API
| Payload ownership | API | What it removes |
| ------------------------------------ | ---------------------------------------------------------------------- | --------------------------------------------------------------------------- |
| One repository owns the store | `preview_collection`, `collect`, `try_collect` | Unreachable records, then unreferenced blobs and chunks. Requires `BlobGc`. |
| Several repositories share the store | `preview_logical_collection`, `collect_logical`, `try_collect_logical` | This repository’s unreachable records only. Requires `MetadataStore`. |
The `casita gc` command uses physical collection for a local repository. Do not add `BlobGc` to a shared backend just to use it: one repository cannot see all physical owners. `CombinedBlobStore` likewise has no `BlobGc` implementation.
## Use logical collection
```rust
let preview = repository.preview_logical_collection().await?;
println!("unreachable records: {}", preview.logical_objects);
let outcome = repository.collect_logical().await?;
if let Some(revision) = outcome.revision {
println!("pruned at repository revision {revision}");
}
```
`collect_logical()` waits for active mutation and retention holds; `try_collect_logical()` returns `Busy` instead. A stale commit fails with a typed revision error, so retry from a new snapshot. Preview and outcome report record counts, not bytes reclaimed: one repository cannot know the physical savings.
An unrooted but committed record still owns its payload. Roots decide which records survive the next logical collection; the shared service tracks payload ownership for every committed record, rooted or not.
## Shared service requirements
The service needs an aggregate owner set, reference count, or equivalent ledger for each physical blob. Its state and ownership changes must share a durable atomic boundary:
| Event | Required service action |
| -------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| A new record commits | Add ownership only if the record was inserted; an idempotent retry must not count twice. |
| Logical collection prunes a record | Remove ownership in the same transaction as the record. |
| A root changes | Leave physical ownership alone; roots affect a later prune. |
| An upload finishes before its record commits | Hold a staging lease so deletion cannot race publication. |
| Ownership reaches zero | Delete only after all write leases end; delayed deletion may leak space but must not remove another owner’s data. |
The state backend still needs revision checks. The payload service also needs writer leases and deletion fencing so a stale deletion request cannot remove a newly published blob. Logical collection alone does not provide these service rules or resolve concurrent changes to names.
## Multi-tenant confidentiality
Logical collection controls lifetime, not access. A hosted service must authorize record and payload reads. Exposing a global `has(BlobId)` or raw blob read would let a tenant test whether another tenant stores known content. Dedupe hits, physical locations, and cross-tenant storage totals can leak the same equality signal.
If equality across tenants must remain private, use separate storage or encryption domains. Removing one tenant’s record does not immediately erase a shared physical blob that another tenant still owns.
## What remains unchanged
`BlobId` still identifies plaintext bytes. Object keys, links, and roots remain logical state in each repository. Authentication, encryption, aggregate ownership, and deletion policy belong to the service around Casita.
See [Garbage Collection](../garbage-collection/) for ordinary collection and [Division of Responsibility](../responsibilities/) for application policy.
# Sync
> How Casita copies verified objects and roots between repositories.
Sync copies selected objects or named roots from one repository to another. `casita sync` uses the library’s `transfer()` engine. The source holds a stable revision while the destination verifies what it receives.
## What gets copied
An object selection copies its forward closure by default. A root selection also copies its full closure, then installs the same name at the destination. For a filesystem root, a path selection copies only the chosen file or directory closure. Casita checks the ancestor directories to establish the path, but does not publish them at the destination.
The receiver skips identical records, rejects conflicting immutable records, and verifies payloads through each object’s format before publication. Peers with compatible chunk stores can avoid sending chunks already present; other payloads use plaintext streaming. Object records publish in bounded batches. Requested roots move only after their complete closures verify, so a failed transfer can leave useful objects behind without moving a root to an incomplete graph. Retrying is safe. Sync does not delete destination roots or objects.
## Complete discovery or faster repeats
By default, Casita examines the full selected source closure, including below objects already present at the destination. `--incremental` stops below an identical object if its destination closure is already verified. This can make repeat transfers faster, but it does not check source descendants below that reused object. A missing or incomplete destination closure is still discovered and verified normally.
In Rust, `TransferOptions::default()` uses complete discovery. Use `with_discovery(TransferDiscovery::ReuseVerified)` for the incremental policy. Discovery queues and visited sets can spill to temporary disk when they exceed the memory threshold. The spill-byte limit is enforced before roots move.
## Source options
Local and S3 repositories can supply source data; the optional SSH adapter reads a remote repository through OpenSSH. An SSH source holds one revision and uses OpenSSH for authentication and host-key checks. It sends whole plaintext payloads over one ordered stream; remote chunk negotiation is not available.
With `--from-blobs`, `--from` supplies the revision, roots, and object records, while a second retained Casita repository supplies payloads. Their revisions may differ because the destination verifies each payload against the selected records. If the payload source lacks required bytes, the transfer fails rather than falling back to `--from`.
See [Synchronize Repositories](../../guides/sync/) for commands and endpoint requirements.
# Verification
> What Casita checks during publication, reads, sync, and integrity scans.
Casita checks stored bytes at different boundaries. The guarantee depends on which operation you use:
| Operation | What is checked |
| ----------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| Import or sync | The destination’s registered format verifier checks each object’s native ID, canonical payload, and exact direct links before publishing its record. |
| Root publication | The complete reachable graph and format relations must verify before the root moves. |
| Chunked read | Each fetched chunk is checked against its `ChunkId` before its bytes are returned. |
| Sequential read from the standard chunked store | An unseeked, complete read checks the `BlobId` at EOF. Stopping early leaves that check unfinished. |
| `open_verified` or `cat --verified` | A Bao proof authenticates bytes against the `BlobId` before each byte is returned; successful EOF also confirms the length. |
A seek or partial read cannot complete that sequential check. Use a verified read when bytes must be authenticated before use. Bao proofs also allow a selected range to be checked against the blob ID without reading the rest of the payload. If the required proof is unavailable, a verified read fails rather than returning unchecked bytes.
`fsck --audit-only` checks the repository’s current snapshot for payload, format, relation, and closure problems without running the repair pass. An unavailable format verifier is reported as `Unchecked`: the data remains retained, but its format validity is not established. See [Errors and Integrity](../../reference/errors/) for findings and [Object Formats](../../reference/object-formats/) for the registered rules.
A hash proves that bytes match a known digest. It does not establish who supplied that digest or who may access the data. Applications need a trusted source for expected root IDs and authorization around services that expose repository content.
# Casitar Archives
> Portable, closure-complete offline interchange for Casita object graphs.
Rust examples on this page use the `experimental` Cargo feature and `casita::experimental`. For the supported built-in workflows, see the [Library guide](/library/).
Casitar moves a complete Casita object graph through a file or stream without a live source repository. The archive names exact root objects and carries their reachable records and plaintext payloads. Import verifies the graph before publishing destination root names.
## Goals and format boundary
Use Casitar when the exchange must preserve Casita’s exact object identities, links, and complete reachable graph. Tar or ZIP is simpler for ordinary files; Git bundles, NAR, CAR, and OCI layouts serve their respective native formats.
Casitar carries a union of the selected roots’ closures. Shared records and payloads appear once. Canonical root and frame ordering, with no timestamps or producer metadata, makes the same graph produce the same archive bytes. Stream limits bound resource use, and the destination chooses its own root names.
CAR does not itself require complete root closures or Casita’s separate namespace-qualified records and shared payloads. Casitar encodes those requirements directly.
## What the archive contains
* A canonical, non-empty set of root `ObjectKey`s.
* Deduplicated plaintext payload frames in strictly ascending `BlobId` order.
* Deduplicated canonical `ObjectRecord` frames in strictly ascending `ObjectKey` order.
* An explicit end marker followed by strict EOF.
It does not copy `casita.sqlite`, physical chunks, compressed representations, bao outboards, repository revisions, or local root names. Those details are local policy or backend layout, not portable graph meaning.
Native Rust callers can use `CasitarReader` and `CasitarWriter` with arbitrary Tokio streams. The codecs check frame lengths, payload hashes, ordering, duplicates, limits, and strict EOF. A successful structural read alone does not prove namespace semantics or closure completeness.
`Repository::export_casitar` accepts root names, exact object keys, or both. It verifies the selected closures under one retention hold, then writes their distinct payloads and records in canonical order. File export stages and syncs a sibling temporary file before atomic publication. The CLI creates a new file by default and requires `--force` to replace one.
## Verification and retention
`Repository::import(CasitarImport::new(...))` hashes every payload, verifies every record against its namespace, and rejects missing or unrelated archive entries. Only a complete, exact closure can publish all destination root mappings together in one revision.
An interrupted import may leave verified, unrooted data for collection, but publishes no selected root. Names must be absent by default. Replacement is conditional on their preflight values remaining unchanged.
## CLI import
```console
$ casita --repository ./source archive create \
--root releases/current --output release.casitar
$ casita archive inspect release.casitar
$ casita archive verify release.casitar
$ casita --repository ./destination archive import release.casitar \
--root releases/imported
```
Create also accepts repeatable exact `--object KEY` selectors and `-` for stdout. Inspect and verify accept `-` for stdin. `inspect` proves only canonical framing, hashes, ordering, strict EOF, counts, and the complete archive digest; it labels that result `structural`. `verify` runs the full importer in an isolated temporary repository and labels success `verified`, without opening the configured durable repository.
Import never invents a destination name. Repeat `--root NAME` exactly once per canonical header root, or select `--root-prefix PREFIX` to map the roots to `PREFIX/0`, `PREFIX/1`, and so on. The default refuses existing names; `--replace` performs compare-and-swap replacement. `--json` emits the stable `casita.archive.v1` report schema. Common maximum-archive, payload, total-payload, payload-count, and record-count flags keep untrusted inputs finite.
## Rust import
Open the reader first so the caller can inspect the archive’s roots and choose destination-owned names before the repository receives any payload. The destination vector must contain one name for every header root, in its canonical order.
```rust
use casita::experimental::{
CasitarReader, CasitarStreamLimits, Repository,
RootName,
};
let repository = Repository::local("./destination").await?;
let input = tokio::fs::File::open("release.casitar").await?;
let reader = CasitarReader::open(input, CasitarStreamLimits::default()).await?;
let destinations = reader
.header()
.roots()
.iter()
.enumerate()
.map(|(index, _)| RootName::try_from(format!("releases/imported/{index}")))
.collect::, _>>()?;
let report = repository
.import(casita::import::CasitarImport::from_reader(reader, destinations))
.await?;
println!("imported {} record(s)", report.records_inserted);
```
# Quick Start
> Install Casita, store a directory, restore it, and copy it to another repository.
Casita requires Rust 1.94.1 or newer. From a source checkout, install the CLI:
```console
$ git clone https://github.com/cachix/casita.git
$ cd casita
$ cargo install --path crates/casita
```
These commands use the repository checkout as the example project. Run them from its top-level directory.
## Store a directory
Create a local repository at `./cache`, then import Casita’s source directory under the name `projects/demo`:
```console
$ casita --repository ./cache init
$ casita --repository ./cache import ./crates/casita/src --root projects/demo
$ casita --repository ./cache root ls
```
The import prints a directory object key, such as `casita.directory.v1:...`. The root keeps that directory and its contents available across restarts.
## Inspect and restore it
Use the key printed by `import` in place of `casita.directory.v1:...`:
```console
$ casita --repository ./cache tree list casita.directory.v1:...
$ casita --repository ./cache checkout casita.directory.v1:... ./restored
```
Checkout requires an absent or empty destination. By default, it also creates an `auto/checkout/...` root that retains the restored graph. Use `--no-root` if you do not need that extra root.
## Copy it to another repository
```console
$ casita sync --from ./cache --to ./mirror --root projects/demo
$ casita --repository ./mirror root ls
```
The destination verifies the graph before setting its root. For SSH sources, path selection, and retry behavior, see [Synchronization](../guides/sync/).
## Check and collect
```console
$ casita --repository ./cache fsck --audit-only
$ casita --repository ./cache gc --dry-run
```
`fsck --audit-only` checks integrity without repairing physical data. The GC dry run shows what collection would remove. To collect, run `gc` without `--dry-run`. Named roots and active operations protect their reachable data.
Without `--repository`, `casita init` attaches a project directory to Casita’s per-user repository using a `.casita` workspace marker. See the [CLI reference](../reference/cli/#repository-and-filesystem-commands) if you prefer that workflow.
# Add a New Importer
> Turn a new input into verified Casita objects, with an optional CLI command.
An importer reads one kind of input and publishes objects through a repository. Start with an existing object format when it fits. A new parser for text, an archive, or a remote feed can store ordinary blobs or canonical directories. If the input needs its own object keys or links that the receiver must verify, [define an object format](../custom-formats/) as well.
The Rust `Importer` trait is available to library users. Implementing it does not add a command to `casita import`; that requires a separate CLI change.
## Define the contract
Before writing data, decide:
* Which input forms are accepted, and how large they may be.
* How paths, links, duplicates, and malformed records are handled.
* Which object format represents the result and which root, if any, retains it.
* What the caller receives after publication and what can remain after failure.
Keep meaning that must survive sync in an `ObjectFormat`. An importer’s parser can reject bad input, but a destination can enforce only the format’s verification rules. For filesystem input, validate paths and link targets before making directory entries. Stream an archive or use handle-rooted access to a source tree rather than extracting untrusted input first.
## Reuse a built-in format
This example accepts bounded, newline-terminated UTF-8 text and stores it as an ordinary blob. It uses the `experimental` feature for Casita’s `async_trait` re-export; the `Importer` trait and `BlobImport` request themselves are part of the supported library API.
```rust
use casita::{
import::{BlobImport, Importer}, ObjectKey, Repository, RootName,
};
use casita::experimental::async_trait;
use std::io;
use tokio::io::{AsyncRead, AsyncReadExt};
struct LineImport {
reader: R,
root: RootName,
}
#[async_trait]
impl Importer for LineImport {
type Report = ObjectKey;
type Error = io::Error;
async fn import(self, repository: &Repository) -> Result {
const MAX_BYTES: u64 = 64 * 1024;
let mut bytes = Vec::new();
self.reader.take(MAX_BYTES + 1).read_to_end(&mut bytes).await?;
if bytes.len() as u64 > MAX_BYTES
|| std::str::from_utf8(&bytes).is_err()
|| !bytes.ends_with(b"\n")
{
return Err(io::Error::new(io::ErrorKind::InvalidData, "invalid line"));
}
let key = repository
.import(BlobImport::new(bytes.as_slice(), self.root))
.await
.map_err(io::Error::other)?;
Ok(key)
}
}
```
`BlobImport` stages the bytes, verifies their blob identity, and publishes the root. The line rule above applies only during this import. A receiver verifies the blob’s bytes and identity, but does not enforce the line rule. Put that rule in a custom `ObjectFormat` if every repository must enforce it.
## Publish a graph
For several linked objects, use the lower level API under `casita::experimental`. Keep one `MutationSession` from the first payload write through root publication. It protects staged bytes from collection.
Stage content with `stage_blob_reader` or `stage_blob`, canonical directories with `stage_directory`, and registered objects with native keys through `stage_object` or `stage_existing`. Each call returns a sealed `StagedObject` after format verification. Publish those values rather than manually constructed `ObjectRecord` values.
Stage children before parents. Publish bounded intermediate batches with `publish_unrooted`, then use `publish_rooted` once the complete closure exists. A failed import may leave collectible unrooted objects, but does not move the root to a partial graph. If replacement depends on an earlier root value, use `publish_if_roots_match`. Use `publish_at_revision` when the entire publication must match one observed repository revision.
## Add a CLI command if needed
To expose an importer in Casita’s CLI, add it to the CLI’s `ImporterKind` list and `casita import` dispatch. Parse the location, limits, and root options, then call the reusable Rust importer. Keep verification and publication in the library.
Give flags for each importer distinct names, such as `--tar-max-entries`. Automatic detection should use a small, bounded structural probe rather than a filename extension. Standard input needs an explicit `-i` unless the CLI can replay the bytes it probed. Document the required Cargo feature, accepted input forms, limits, root behavior, output, and failures. A library-only importer needs no CLI command.
Test valid and malformed input, each size limit, unsafe paths or links when relevant, conflicting entries, object identity and links, interrupted input, and final root publication. Test CLI parsing separately if you add a command.
See [Import Semantics](../../concepts/imports/) for the shared lifecycle and the [Experimental Rust API](../../reference/experimental-rust-api/#publication) for staging and publication methods.
# Import a Casitar Archive
> Restore a portable, verified Casitar closure archive through the generic importer interface.
Rust examples on this page use the `experimental` Cargo feature and `casita::experimental`. For the supported built-in workflows, see the [Library guide](/library/).
Casitar archives carry a complete, canonical closure: payloads, immutable records, and one or more source root declarations. Importing verifies the entire stream before atomically publishing destination-owned root names. It does not recreate the source repository namespace.
## CLI
`import` detects a Casitar file from its header. Select `-i casitar` for standard input. Supply an explicit destination mapping: repeat `--casitar-root` once for every archive root, in the canonical header order, or use `--casitar-root-prefix` to map them to `PREFIX/0`, `PREFIX/1`, and so on.
```console
$ casita --repository ./cache import release.casitar \
--casitar-root releases/current
$ casita --repository ./cache import -i casitar - \
--casitar-root-prefix releases/2026-08
```
Destination names must be absent by default. `--casitar-replace` replaces them only when their preflight values have not changed before publication. The common `--root` option is intentionally not accepted for this importer.
The importer accepts `-` for standard input and uses finite stream limits. Use `--casitar-max-archive-bytes`, `--casitar-max-payload-bytes`, `--casitar-max-total-payload-bytes`, `--casitar-max-payloads`, and `--casitar-max-records` to tighten or override them for a deployment.
`archive import` remains available for its JSON report and its historical archive-specific flags. `import -i casitar` is the common importer command and uses only `--casitar-*` options.
## Rust
Open the bounded stream, choose destination names, and pass the reusable `CasitarImport` value through `Importer`. The destination vector is in the same order as the roots in the archive header.
```rust
use casita::import::{CasitarImport, Importer as _};
use casita::experimental::{
CasitarReader, CasitarRootConflictPolicy, CasitarStreamLimits, Repository, RootName,
};
let repository = Repository::local("./cache").await?;
let input = tokio::fs::File::open("release.casitar").await?;
let reader = CasitarReader::open(input, CasitarStreamLimits::default()).await?;
let report = CasitarImport::from_reader(reader, [RootName::try_from("releases/current")?])
.with_conflict_policy(CasitarRootConflictPolicy::RequireAbsent)
.import(&repository)
.await?;
println!("published at {}", report.destination_revision);
```
Use `ReplaceIfUnchanged` only when replacing a root is intentional. A failed import never publishes any mapped destination root, though verified unrooted data may remain available for collection.
See [Portable Casitar Archives](../../design/casitar/) for the wire format and [Import Semantics](../../concepts/imports/) for the shared publication model.
# Define a Custom Object Format
> Verify a new object's identity, payload, and links in every repository that uses it.
Use a custom `ObjectFormat` when an object needs its own key namespace or format rules that another repository must enforce. An importer can parse a source, but the format verifier decides whether its stored object is valid. The destination must register the same verifier before it can accept those objects through sync.
This API is in `casita::experimental` and requires the `experimental` Cargo feature. For ordinary blobs and filesystem trees, start with the [supported library API](/library/) and [importers](../adding-an-importer/).
## Verify an object
An `ObjectFormat` owns one namespace. Its `verify` method checks the native ID, reads the full payload within a bound, validates its encoding, and returns the exact direct links. The repository supplies the read context and creates a sealed `VerifiedObject`; a caller cannot publish an unchecked record in its place.
This example stores newline-terminated UTF-8 documents. Each key uses the BLAKE3 digest of the complete document as its native ID. Documents have no links. The example registers the format, stages one document, and names it:
```rust
use std::sync::Arc;
use casita::experimental::{
async_trait, BlobFormat, Digest, DirectoryFormat, FormatError, FormatLimits,
FormatRegistry, MemoryBlobStore, MemoryMetadataStore, NamespaceId, ObjectFormat,
ObjectKey, Repository, RootName, VerificationContext, VerifiedObject,
};
struct DocumentFormat {
namespace: NamespaceId,
}
#[async_trait]
impl ObjectFormat for DocumentFormat {
fn namespace(&self) -> &NamespaceId {
&self.namespace
}
async fn verify(
&self,
mut context: VerificationContext<'_>,
limits: &FormatLimits,
) -> Result {
let expected = context.key().native_digest().ok_or_else(|| {
FormatError::NativeIdLength {
namespace: self.namespace.clone(),
actual: context.key().native_id().len(),
}
})?;
let bytes = context
.read_to_end_bounded(
(1024 * 1024).min(limits.max_metadata_bytes).min(limits.max_payload_bytes)
)
.await?;
let text = std::str::from_utf8(&bytes).map_err(|error| {
FormatError::InvalidPayload {
namespace: self.namespace.clone(),
message: error.to_string(),
}
})?;
if !text.ends_with('\n') {
return Err(FormatError::InvalidPayload {
namespace: self.namespace.clone(),
message: "document must end with a newline".into(),
});
}
let actual = context.observed_digest();
if actual != expected {
return Err(FormatError::NativeIdentityMismatch {
key: context.key().clone(),
expected,
actual,
});
}
context.finish(Vec::new())
}
}
let namespace: NamespaceId = "example.document.v1".parse()?;
let formats = FormatRegistry::new([
Arc::new(BlobFormat::default()) as Arc,
Arc::new(DirectoryFormat::default()) as Arc,
Arc::new(DocumentFormat { namespace: namespace.clone() }) as Arc,
])?;
let repository = Repository::with_formats(
MemoryBlobStore::new(),
MemoryMetadataStore::new()?,
formats,
FormatLimits::default(),
);
let bytes = b"hello\n";
let key = ObjectKey::new(namespace, Digest::hash(bytes).as_bytes().to_vec())?;
let mutation = repository.mutation_session().await?;
let staged = mutation.stage_object(key.clone(), bytes).await?;
mutation
.publish_rooted(vec![staged], RootName::try_from("documents/hello")?, key)
.await?;
```
`read_to_end_bounded` consumes the whole payload or fails at the configured limit. `context.finish` requires end of input, records the payload digest and size, and seals the declared links. In this example, an empty link list is part of the format’s meaning.
## Register every format you need
`FormatRegistry::new` contains exactly the formats you list and rejects two verifiers for the same namespace. The example includes raw blobs, canonical directories, and documents. Add any other formats your repository must read or publish. The registry is fixed when the repository opens, so repositories that exchange custom objects must each register the matching verifier.
For a format with child objects, return their canonical direct keys from `context.finish(links)`. Those links control closure checks, roots, sync, and collection. Override `verify_links` only when validity also depends on a declared direct child’s verified record or payload. Read it through `DirectLinkView`; keep verification deterministic and independent of the network, clock, and mutable roots.
## Test the format
Test valid input, malformed or truncated payloads, wrong native IDs, size and link limits, link order, and relations to missing or incompatible children. Check that publication rejects an invalid object and that a receiving repository with the same registry verifies it after sync.
See [Object Formats](../../reference/object-formats/) for built-in namespaces and the [Experimental Rust API](../../reference/experimental-rust-api/#object-formats) for the verifier contract.
# Capture and Restore a Filesystem Tree
> Import, inspect, retain, refresh, and safely materialize a canonical filesystem graph.
This workflow turns one directory into a verified immutable graph, gives it a durable name, and restores the exact graph later.
## Import and retain the tree
### CLI
```console
$ casita --repository ./cache import ./project --root projects/demo
```
Directories select the filesystem importer automatically. The command prints the resulting `casita.directory.v1` key. The named root retains that directory and every object reachable through its verified links.
Imports normally reuse an earlier file result when device, inode, size, and timestamps are unchanged. Force every file to be reread when that assumption does not suit the source:
```console
$ casita --repository ./cache import ./project \
--root projects/demo --filesystem-rehash
```
### Rust
`FilesystemImport` publishes the tree and its root in one operation. Use `FilesystemImport::new(...).reread(true)` when every regular file must be read and hashed again.
```rust
use casita::{Repository, RootName};
let repository = Repository::local("./cache").await?;
let root = RootName::try_from("projects/demo")?;
let tree = repository.import(casita::import::FilesystemImport::new("./project", root)).await?;
println!("{tree}");
```
## Inspect the graph
```console
$ casita --repository ./cache root ls projects
$ casita --repository ./cache object show casita.directory.v1:...
$ casita --repository ./cache tree list casita.directory.v1:...
$ casita --repository ./cache cat casita.blob.v1:...
```
`object show` exposes the logical key, physical payload, exact forward links, and closure status. `tree list` interprets a canonical directory; `cat` writes a blob’s bytes to standard output.
## Restore safely
```console
$ casita --repository ./cache checkout casita.directory.v1:... ./restored
```
The destination must be absent or empty. Checkout first builds a sibling staging directory on the destination filesystem, then renames it into place. It either recreates the graph’s exact names or fails without a partial checkout; this catches case-folding, Unicode-normalization, and native-name conflicts on the actual target filesystem. Checkout uses handle-relative writes so a path component swapped for a symlink during materialization cannot redirect writes outside the destination. Stored symlinks are recreated rather than followed.
Successful checkout creates an `auto/checkout/...` root by default. Use `--no-root` only when another root already retains the graph or the restored copy is intentionally disposable.
## Release data deliberately
```console
$ casita --repository ./cache root rm projects/demo
$ casita --repository ./cache gc --dry-run
$ casita --repository ./cache gc
$ casita --repository ./cache fsck
```
Removing a name only makes its unshared closure eligible for collection. The dry run shows what a real collection would remove.
Read [Imports](../../concepts/imports/) for cache assumptions, [Roots and Retention](../../concepts/roots-and-retention/) for liveness, and the [CLI Reference](../../reference/cli/) for exact syntax.
# Preserve and Serve a Native Git View
> Import native Git objects without translating their identity, then inspect, materialize, or serve an immutable ref view.
Casita preserves Git’s native object identity and stores an immutable view of selected refs. It is a retention and read path, not a Git authoring or push service.
## 1. Import a repository
### CLI
```console
$ casita import ./project \
--git-view upstream \
--git-ref refs/heads/main
```
`SOURCE` may be a working tree or bare repository; Casita detects both from their metadata. Without explicit `--git-ref` arguments, Casita selects local branches and tags. When `--git-view` is omitted during auto-detection, Casita uses the source directory’s basename. Import verifies native Git objects, publishes one immutable ref view, and sets the root `git/upstream`.
`--git-concurrency` bounds active object staging (default 16), and `--git-max-buffered-bytes` bounds decoded source bytes held by staging futures (default 67108864, or 64 MiB). Both must be positive. Set concurrency to 1 for serial staging. An object larger than the byte budget runs alone. Gix caches, delta-decoding workspace and payload-store buffers are outside this budget. Source decoding remains synchronous; object verification and storage overlap.
### Rust
Enable Casita’s `git` feature, then pass a local working tree or bare repository and the selected refs. An empty `refs` list selects local branches and tags. `max_cached_pack_bytes` defaults to 8 GiB and retains a verified exact source pack for fast full clones; set it to zero when minimizing stored bytes is more important than full-clone throughput.
`GitImport::with_concurrency` and `with_max_buffered_bytes` configure the same limits using nonzero integers. `NativeGitImportOptions` also exposes `concurrency` and `max_buffered_bytes` fields.
```toml
[dependencies]
casita = { git = "https://github.com/cachix/casita", default-features = false, features = ["git"] }
```
```rust
use casita::{import::GitImport, Repository};
let repository = Repository::local("./cache").await?;
let outcome = repository
.import(
GitImport::new("./project", "upstream")
.with_refs(["refs/heads/main"])?,
)
.await?;
println!("{} objects in {}", outcome.objects, outcome.view);
```
## 2. Inspect the view
```console
$ casita --repository ./cache git show upstream
```
The output includes the view key, object format, optional default ref, and direct or symbolic refs. A view captures one exact ref state; importing again creates a new immutable view and repoints the destination-owned root.
## 3. Materialize an exact tree
```console
$ casita git checkout \
git.sha1.tree.v1:... ./tree
```
The target must be absent or empty. Gitlinks fail by default because Casita does not silently fetch or materialize submodules; `--skip-gitlinks` creates empty directories for them instead.
## 4. Offer read-only clones
With smart-HTTP support enabled, bind one immutable view:
```console
$ casita git serve upstream \
--listen 127.0.0.1:9418
```
The printed URL supports read-only full and shallow fetches. The service does not implement receive-pack, branch mutation, review, merge, or repository administration. Put authentication, TLS, network exposure, and process supervision around it as deployment policy requires.
See [Object Formats](../../reference/object-formats/) for Git namespaces and the [CLI Reference](../../reference/cli/#native-git) for exact command forms.
# Import an OCI Image
> Import a container image, keep its original layers, and check out its merged filesystem.
Import a container image from a registry into Casita. You can keep the image’s original layers and also combine them into a filesystem to check out locally.
## 1. Enable OCI support
From a Casita source checkout, install the CLI with the `oci` feature:
```console
$ cargo install --path crates/casita --features oci
```
See [Cargo Features](../../reference/cargo-features/) for other build options.
## 2. Import an image
This example imports Alpine and keeps both the image and its filesystem in the repository at `./cache`:
```console
$ casita --repository ./cache import \
-i oci docker.io/library/alpine:latest \
--root images/alpine \
--oci-rootfs-root filesystems/alpine
```
The two names refer to different things:
| Option | What Casita keeps under that name |
| -------------------------------------- | --------------------------------------------------------------------------------- |
| `--root images/alpine` | The image layout: its manifest, config, and original layer archives. |
| `--oci-rootfs-root filesystems/alpine` | The merged filesystem: the files and directories left after applying every layer. |
The names must differ. To keep only the image layout, omit `--oci-rootfs-root`.
Casita verifies the downloaded content before publishing either name. If the import fails, any previous image or filesystem under those names remains available. An interrupted import may leave unused objects that garbage collection can remove.
## 3. Check out the filesystem
The import prints a `root` key for the image layout and a `rootfs` key for the merged filesystem. Copy the directory key printed after `rootfs`, then use it in place of `ROOTFS_KEY` below:
```console
$ casita --repository ./cache checkout ROOTFS_KEY ./alpine-rootfs
```
Use the same repository path as the import. The destination must be absent or empty. The merged filesystem can also be mounted through `casita-fs`.
To check out the original image layout instead, use the key printed after `root`. That directory contains `oci-layout`, `index.json`, and `blobs/sha256/`.
## Choose a platform
For an image that offers several platforms, Casita selects your current operating system and architecture. Add `--oci-platform` to select another:
```console
$ casita --repository ./cache import \
-i oci docker.io/library/alpine:latest \
--root images/alpine-arm64 \
--oci-rootfs-root filesystems/alpine-arm64 \
--oci-platform linux/arm64
```
Each import keeps one selected platform. Other platforms from the image index are not included in its layout.
The CLI connects over HTTPS and uses anonymous registry access. Use `--oci-http` for a registry served over plain HTTP. Private registry credentials can be supplied through the Rust API below.
## How layers become a filesystem
Casita applies layers in image order. Later layers can replace files or remove them using deletion markers called *whiteouts*. An opaque directory marker removes children inherited from earlier layers while keeping new children from its own layer.
The importer supports uncompressed tar, gzip, and zstd layers, including Docker gzip layers. It preserves file contents, directory names, symbolic link targets, and the executable bit. Hardlinks share the same stored file content, and symbolic links are stored without following their targets.
Owners, timestamps, other permission bits, extended attributes, and hardlink inode identity are not preserved in the merged filesystem. The original layer archives retain this metadata. Sparse files and special files such as devices and FIFOs are rejected when building a merged filesystem. Windows checkout refuses absolute symbolic link targets.
Each layer streams into storage as it downloads. When a merged filesystem is requested, the same transfer also feeds the decoder and stores the files inside it. No complete layer is buffered in memory or reread from storage. A repeated layer descriptor is downloaded once for each occurrence.
Casita checks each downloaded blob’s digest and size. It also checks the uncompressed layer digest, called a *DiffID*, against the image config. Both named roots are updated together only after all checks pass.
## Resource limits
Large file contents stream through the importer, while directory metadata stays in memory. The following default limits bound an import:
| Resource | Default limit |
| ----------------------------------------- | ---------------------------------- |
| Manifest or config JSON | 16 MiB each |
| Layers | 1024 |
| One downloaded config or layer blob | 256 GiB |
| All downloaded config and layer blobs | 1 TiB |
| Decoded tar bytes | 1 TiB per layer and in total |
| Layer entries and merged filesystem nodes | 1,000,000 each |
| Pathname or hardlink target | 4096 bytes |
| One regular file | 256 GiB |
| File content stored across all layers | 1 TiB, including overwritten files |
Use `--oci-max-blob-bytes` and `--oci-max-total-blob-bytes` to change the download limits. For a merged filesystem, `--oci-rootfs-max-bytes` sets the decoded tar limit per layer and in total, and `--oci-rootfs-max-entries` sets both entry limits. Byte limits are supplied as integers.
The registry client buffers manifests before Casita checks their JSON size. Building a merged filesystem also buffers the config within its limit so Casita can read the expected layer digests.
## Use from Rust
Enable the `oci` feature and pass an `OciImport` request to `Repository::import`. `OciImport::new` takes an image reference and the name for its layout. Add `with_rootfs` with a different name to request a merged filesystem; its key is returned in `OciImportReport::rootfs`.
Use `with_auth` to supply registry credentials, `with_platform` to select a platform, and `with_client` to configure the registry client, including TLS. `with_limits` accepts `OciImportLimits`, and `with_rootfs_limits` accepts `OciRootfsLimits` for the remaining resource limits.
See the [CLI Reference](../../reference/cli/#import--i-oci) for the full command syntax and the [Library guide](../../library/) for repository setup.
# Operate a Local Repository
> Check roots and integrity, collect unused data, and back up a coherent repository.
Roots tell Casita which graphs to keep. Routine operation is to inspect those roots, preview collection, and audit integrity before making changes.
## Routine checks
```console
$ casita --repository /var/lib/casita root ls
$ casita --repository /var/lib/casita gc --dry-run
$ casita --repository /var/lib/casita fsck --audit-only
```
`gc --dry-run` computes the collection plan without deleting data. `fsck --audit-only` checks physical payloads, logical records, and reachable closures without running the repair pass. It can report collectible residue or an unavailable format verifier without classifying the repository as corrupt. Read [Errors and Integrity](../../reference/errors/) for the report categories.
## Collect unreachable data
```console
$ casita --repository /var/lib/casita gc
```
Collection removes unreachable records, then unreferenced payloads and chunks. Active mutations and retained readers pin the data they need while unrelated garbage remains collectible. The local profile can also try collection when a mutation starts on a filesystem at least 80% full. It may then release roots explicitly marked evictable, preserving permanent roots and data held by active operations. Explicit capacity monitoring and collection remain useful because busy work can defer that attempt.
## Back up and restore
Back up the repository directory as one unit:
1. Stop or quiesce every process that can write, retain, or collect.
2. Copy the complete directory, including `casita.sqlite` and `blobs`.
3. Restore that complete copy, then open it and run `fsck` before relying on it.
Ordinary file copy tools do not participate in Casita’s locks. A live file-by-file copy may combine incompatible state and payload revisions. Treat the restored copy as a separate repository; its revision tokens cannot order changes against the original.
## Respond to a failed integrity check
Preserve the affected repository before attempting recovery. A normal `fsck` may rebuild a corrupt Bao outboard from verified local bytes. With an independently verified local replica, `--source` can replace a missing or corrupt physical representation:
```console
$ casita --repository /var/lib/casita fsck --dry-run
$ casita --repository /var/lib/casita fsck --source /srv/casita-replica
```
`--dry-run` previews physical repair actions without writing. Repair never invents object records or changes roots. For reachable missing or invalid content that cannot be repaired from a trusted replica, restore a coherent backup or re-import and re-sync. For collectible residue, run `gc` and audit again. For an `Unchecked` namespace, use a build with its format verifier.
Read [Local Repository](../../reference/local-repository/) for layout and [Garbage Collection](../../concepts/garbage-collection/) for collection order.
# Run Applications from Casita
> Import an application tree, select its executable, and run a retained copy.
`casita run` takes an existing named directory root, checks out its tree into a private temporary directory, and starts an executable from that tree. It does not build the application or fetch a missing root.
For a root already published by a build tool, the command can be as short as:
```console
$ casita run cargo/builds/uv -- --version
```
## Import and run a tree
Install the CLI from a Casita checkout or follow the [Quick Start](../../getting-started/). This POSIX example makes one executable:
```sh
mkdir -p output/bin
cat > output/bin/hello <<'SH'
#!/bin/sh
printf 'Hello from Casita\n'
SH
chmod +x output/bin/hello
```
```console
$ casita --repository ./cache import ./output --root apps/hello
$ casita --repository ./cache run apps/hello
Hello from Casita
```
Use the same repository for import and run. A compiled application can include its required libraries and assets in `output/`. The program must also be compatible with the host’s operating system, architecture, and installed runtime dependencies.
## Choose an executable
Casita searches the tree recursively. On Unix, files need an executable mode bit; on Windows, candidates need `.exe` or `.com`. If several candidates exist, `run` lists them and asks for `--bin`:
```console
$ casita --repository ./cache run apps/hello --bin hello
$ casita --repository ./cache run apps/hello --bin release/hello
```
A filename must match uniquely. A relative path selects one exact candidate; use `--bin ./hello` to distinguish a root-level file. On Windows, `--bin hello` can match `hello.exe` or `hello.com`. Stored symlinks to executables inside the tree can be selected, but directory links and links outside the tree are not searched. Casita does not search your shell’s `$PATH` for a candidate.
Put application arguments after `--`. The child inherits the current working directory, environment, and standard streams. Its exit status becomes Casita’s exit status.
## Shorten root names in a project
Run `casita init` in the project directory, then add these settings to its `.casita` marker without removing the marker header or workspace UUID:
```toml
[run]
default-scope = "cargo"
[run.scopes]
cargo = "cargo/builds"
go = "go/builds"
```
```console
$ casita run cargo:uv -- --version
$ casita run go:server -- --port 8080
$ casita run uv -- --version
```
The last command expands to `cargo/builds/uv` through the default scope. A full root path stays literal, and `/uv` bypasses scope expansion to select the literal root `uv`. Casita finds the closest `.casita` marker in the current directory or its ancestors. `run` does not add the workspace UUID to root names. Unknown scopes fail instead of selecting another root.
## Copy or update an application
Sync a named root before running it from another local repository:
```console
$ casita sync --from ./cache --to ./mirror --root apps/hello
$ casita --repository ./mirror run apps/hello
```
Import a new output under the same root to update later runs. A process already running keeps the exact tree it selected, even if that root changes or collection runs. Normal completion removes the temporary checkout and releases its retention hold. The child runs with your permissions; `run` is not an execution sandbox.
See the [run command reference](../../reference/cli/#run) for exact selection, signal, and cleanup behavior.
# Maintain an S3 Repository
> Inspect holds, shut down cleanly, and recover interrupted collection.
Every runner sharing an S3 repository must use the same bucket and prefix and support its WAL3 admission protocol. Stop older binaries before upgrading because they do not observe the current holds. The S3 profile requires the `s3` feature. Recovery APIs below are in `casita::experimental`.
## Routine checks
```sh
casita holds s3://my-bucket/repository
casita holds s3://my-bucket/repository --json
```
`casita holds` reads collector ownership and the state and coordination pin ledgers without waiting for repository admission. It uses standard AWS credentials. Each ledger is a separate snapshot, so use exact tokens rather than matching entries by writer name or time. Inspection never releases a hold. A listed token does not prove its owner has stopped.
For shared root ownership, see [Share an S3 Repository Across Owners](../s3-multi-owner/).
## Collection and reads
Mutation sessions and retained reads create durable pins. They protect staging data, selected payloads, snapshots, and metadata files while collection runs. Keep a retention hold until its payload stream closes. A WAL3 cursor alone does not protect physical payloads.
Only one collector runs at a time. `collect` waits for ownership; `try_collect`, `try_collect_logical`, and `try_vacuum` return `Busy` when another collector or a conflicting pin update prevents a pass. Holds do not expire. An abandoned collector token blocks later `collect` calls until it is recovered.
Collection first prunes unreachable logical records, then updates the payload catalog and deletes retired packs and manifests. An interrupted pass may leave physical garbage; `vacuum` reclaims deferred packs and obsolete catalog files. Rooted data and active pins remain protected.
## Clean shutdown
For a library application, stop new work, await running operations, drop sessions and readers, then await `casita::experimental::flush_repository_leases()` before stopping Tokio. Dropping a session schedules its durable release; the final flush waits for those releases and background catalog work. A failed release reports an error and leaves the durable pin in place. The CLI drains releases on exit.
## When a command is blocked
Run `casita holds s3://BUCKET/PREFIX --json` to see every token, pin scope, released history, prune fence, and deletion claim. Admission warnings show a short list of blocking tokens by default. For acquisition and release events, use `--log-filter casita=info`; `casita=debug` adds pin events.
Writer names are diagnostic. The CLI uses `CASITA_WRITER` when set, otherwise a process ID and random suffix. Even an explicit name can be reused. Correlate the exact token with process logs before deciding it is abandoned.
## Recover an abandoned hold
Holds have no TTL. A paused process can still resume and read or publish data, so elapsed time is never sufficient evidence that its hold is abandoned.
1. Terminate the owning runner and prevent it from resuming. For an interrupted exclusive collector, stop every runner and allow outstanding backend requests to settle before recovery.
2. Run `casita holds s3://BUCKET/PREFIX --json` to inspect operational tokens and both pin ledgers. For interrupted payload collection, record the exact `state.collector` token as well as the operational collector token. The two tokens identify different ownership records.
3. Call `release_abandoned_repository_hold(&hold.token)` for the exact abandoned token. Reusing a diagnostic writer name does not grant ownership; replaying recovery for an old token cannot remove a newer token.
4. For interrupted payload collection, call `Repository::recover_s3_collection(bucket, prefix, writer, &collector_token)`. It obtains collector ownership before opening payload discovery and can therefore reopen through an abandoned logical-prune fence. It preserves live pins and inherited deletion claims throughout marking and metadata commit, then finishes the sweep and catalog publication. For an already-open generic repository, use `recover_collection(&collector_token)`.
5. Drain releases, check the remaining inventories, and run an integrity audit before resuming traffic. If a recovery attempt fails, inspect again: its new collector token may have replaced the abandoned token while preserving the original fence and claims. Never substitute an older token for the current one.
The operational collector token and `state.collector` are different records. Releasing the operational hold does not clear a prune fence or deletion claim.
For other abandoned records, first establish that the owner and its outstanding requests have stopped:
| Record | Recovery API |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------- |
| Reader or writer pin | `release(&exact_pin_token)` on `state.pin_store()` or `state.repository_coordination_pin_store()`, according to its ledger |
| State WAL shard claims | `state.recover_wal_deletions(exact_claim_tokens)` |
| Coordination WAL claims | `state.recover_repository_coordination_deletions(exact_claim_tokens)` |
WAL claim recovery needs the complete current token set for that ledger. It verifies that claimed paths are unreferenced and retains claims through retries. Pin release clears only the named pin. Reinspect after a failure before retrying because the current tokens may have changed.
## Maintain the WALs
`Wal3MetadataStore::collect_wal(reader_grace_period)` collects obsolete state-log shards. `collect_repository_coordination(reader_grace_period)` compacts the separate admission log without expiring holds. Both acquire collector ownership and can run alongside ordinary readers and writers. Choose a grace period longer than the longest manifest-to-fragment metadata read. Pins protect files in use; unsettled deletion claims prevent path reuse until recovery. A failed collection retains its token for recovery.
# Share an S3 Repository Across Owners
> Use separate root names and one collector for several applications.
Several applications can share one S3 repository and deduplicate identical content. Give each owner a separate root prefix, such as `owner-a/current` and `owner-b/releases/42`. Casita retains the union of their named graphs and active reads; one collector removes data nobody needs.
Root ownership, permissions, scheduling, and peer discovery are application decisions. This guide uses the supported `casita::Repository` API except where it explicitly names an experimental type. The [`s3_multi_owner` integration tests](/library/#test-remote-application-workflows) exercise separate writer, reader, and collector processes.
## The composition
Every process opens `Repository::s3(bucket, prefix, writer)` with the same bucket and prefix. The writer name is diagnostic; random durable tokens own pins and collector passes. Casita treats root prefixes as opaque, so the application must control every writer for each owner’s prefix.
Owners may point to the same object key or share chunks through different graphs. Removing one root leaves content retained by another root or active pin. Any process may collect, but only one collector runs at a time.
## What retention is and is not
Named roots retain their complete graphs. Online pins protect active work:
| Read API | Retention scope |
| ----------------------- | ----------------------------------------------------------------------------------------------------------------- |
| `retained_reader` | The whole snapshot at its generation, including data from other owners. It can still resolve names removed later. |
| `open`, `open_verified` | Only the selected object’s closure. Unrelated data stays collectible. |
Pins survive root removal and dropping the repository handle. Process death does not expire them. A long lived retained reader may delay reclamation of older garbage from other owners; use closure-scoped opens for long reads when possible.
Retention does not grant access control. Any client with credentials for the shared prefix can read its objects. Use separate repositories or bucket policy when owners need isolation.
## Conditional roots compare current values
`compare_and_set_root`, `remove_root`, and `commit` checks compare current targets, not a name’s history. After removing a name, a delayed create that expects absence could set its old target again.
To reject that replay, advance an owner-scoped fence root in the same commit as each shared-name change. Check the fence value observed before the change:
```rust
use casita::{MetadataChange, MetadataCheck, MetadataCommitResult};
let result = repository
.commit(
vec![
MetadataCheck::Root { name: fence.clone(), expected: Some(epoch_0.clone()) },
MetadataCheck::Root { name: shared.clone(), expected: None },
],
vec![
MetadataChange::SetRoot { name: shared.clone(), target: key.clone() },
MetadataChange::SetRoot { name: fence.clone(), target: epoch_1.clone() },
],
)
.await?;
assert!(matches!(result, MetadataCommitResult::Committed { .. }));
```
A later removal checks `epoch_1` and advances to `epoch_2`. A delayed duplicate create still expects `epoch_0`, so it conflicts even if the shared name is absent again. Put the fence check first: conflicts report the first failing check in caller order. Fence values are ordinary immutable objects. The application decides its claim IDs and replay policy.
## Publication, collection and retries
Collection marks one repository revision. Overlapping publication can return `StaleRevision`; a racing pin update can return `Busy`. Both are retryable without a logical prune. Schedule collection when publications leave a quiet window. `collect` waits for collector ownership; schedulers can use `try_collect` to receive `Busy` instead.
An integrity audit may report collectible unrooted objects. Use `IntegrityReport::is_healthy` to distinguish those findings from reachable corruption; require `is_clean` after a completed collection.
## Process failures and recovery
An interrupted reader can leave a durable pin, and an interrupted collector can retain its ownership token and deletion claims. They do not expire. Once the owner and its outstanding requests have stopped, recover by exact token. A release for one owner’s pin cannot remove another owner’s protection.
Follow [S3 maintenance](../s3-maintenance/#recover-an-abandoned-hold) for the complete recovery procedure. These inspection and recovery APIs use `casita::experimental`.
## Selected root readers
A `retained_reader` protects the entire snapshot, so keep it short in a shared repository. Resolving a root and then calling a closure-scoped `open` takes two steps. If the root is removed and collected between them, `open` returns `None`; retry the lookup. The experimental `TransferSelection::Selected` can resolve and pin selected root closures for a transfer session.
# Synchronize Repositories
> Copy objects, roots, or one filesystem path into a verified destination.
Use `casita sync` to copy selected data into another Casita repository. The source holds one stable revision; the destination verifies incoming objects. Choose a named root to keep a complete graph under the same name, or an object key to copy a graph without naming it.
## Choose what to copy
**A named root** copies its full closure and installs the same name at the destination:
```console
$ casita sync --from ./cache --to ./mirror --root projects/demo
```
**An object key** copies its forward closure by default. Repeat `--object` for more keys. `--shallow` copies only the selected objects’ records and payloads; it never makes a root transfer shallow.
```console
$ casita sync --from ./cache --to ./mirror \
--object casita.directory.v1:...
$ casita sync --from ./cache --to ./mirror \
--object casita.blob.v1:... --shallow
```
**One path below a filesystem root** copies just that file or directory closure. `--path` requires exactly one `--root` and cannot be combined with `--object` or `--shallow`:
```console
$ casita sync --from ./cache --to ./mirror \
--root projects/demo --path lib/python3.12 \
--destination-root partial/python3.12
```
Casita verifies the path’s ancestor directories but does not store them at the destination. `--destination-root` gives the selected closure a durable name. Without it, the copied objects are unrooted and may be collected. A symlink encountered as the selected path is reported, but has no standalone object key to copy or root.
## Select endpoints
| Endpoint | Source | Destination | Requirement |
| ---------------------------------------- | ------ | ----------- | ------------------------------------------------------------ |
| Local repository path | Yes | Yes | Standard CLI |
| `s3://BUCKET/PREFIX` | Yes | Yes | `s3` feature |
| `ssh://[user@]host[:port]/absolute/path` | Yes | No | `ssh` feature on both machines and remote `casita` on `PATH` |
For an SSH source, install the binary with SSH support on both machines:
```console
$ cargo install --path crates/casita --features ssh
$ casita sync \
--from ssh://alice@example.com/var/lib/casita \
--to ./mirror --root projects/demo
```
OpenSSH handles authentication, host-key checks, and proxy settings. Casita still verifies the received objects locally. `--writer NAME` sets the diagnostic WAL writer name when an S3 endpoint is involved.
Casita passes `ConnectTimeout=30`, `ServerAliveInterval=15`, and `ServerAliveCountMax=3` to `ssh`, so an unreachable host fails within 30 seconds, and a host or network that stops answering mid-transfer fails the sync within about 45 seconds instead of blocking it. These command-line options take precedence over the same settings in `ssh_config`.
## Read payloads from another repository
Use `--from-blobs` when roots and object records live at one endpoint but the required payloads are available from another:
```console
$ casita sync \
--from ssh://alice@example.com/var/lib/casita \
--from-blobs s3://casita-mirror/releases \
--to ./mirror --root projects/demo
```
This example needs the `cli`, `ssh`, and `s3` features. The payload source must be a Casita repository, not a bucket of arbitrary files. Both source sessions stay retained during the transfer; their revisions may differ. The destination checks payloads against the records from `--from`. A missing payload fails the transfer without falling back to `--from`. An SSH payload source also needs the corresponding object records to serve payloads by object key. If `--from-blobs` is omitted, `--from` supplies both records and payloads.
## Repeat or check a transfer
A failed transfer may leave verified objects in the destination, but requested roots move only after their complete closures verify. Retry the same command to reuse work already done. Sync never removes destination data.
Add `--incremental` for faster repeated transfers when reusing a verified destination closure is sufficient. The default examines the entire selected source closure, including descendants of objects already at the destination. With `--incremental`, those reused descendants are not audited at the source.
To inspect a copied root and run an integrity check:
```console
$ casita --repository ./mirror root ls projects/demo
$ casita --repository ./mirror fsck
```
See [Sync](../../concepts/sync/) for the transfer guarantees and the [CLI Reference](../../reference/cli/#synchronization) for every option.
# Import a Tar Archive
> Stream a tar archive directly into a verified canonical filesystem graph.
The tar importer creates the same canonical filesystem object model as a filesystem import, without extracting untrusted paths to disk first. It accepts regular files, directories, symlinks, hard links, and old-GNU sparse files. PAX GNU sparse members are rejected because their logical expansion is not yet available in the underlying async tar reader.
## CLI
`import` detects a tar file from its first valid header; `import -i tar` also accepts `-` for standard input. The stream must already be decompressed; for example, use a decompressor pipeline for a `.tar.gz` input. The named root is installed only after the whole archive is validated.
```console
$ casita --repository ./cache import release.tar --root releases/current
$ gzip -cd release.tar.gz | casita --repository ./cache import -i tar - \
--root releases/current
```
The command defaults to the same finite archive, entry, pathname, file-byte, total-file-byte, and sparse-expansion limits as `TarImportLimits`. Override them with `--tar-max-archive-bytes`, `--tar-max-entries`, `--tar-max-path-bytes`, `--tar-max-file-bytes`, `--tar-max-total-file-bytes`, and `--tar-max-sparse-expansion-bytes` for a stricter deployment policy.
File finalization is pipelined while the next tar body is read. At most 16 files are being copied, queued, finalized, or verified at once. Set `--tar-max-in-flight-files 1` for serial processing, or choose another positive limit to control the number of open writers. Archive parsing stays sequential; the importer does not buffer whole files. Compression and chunk hashing use the payload backend’s existing workers.
## Rust
`TarImport` accepts an already-decompressed asynchronous tar stream, so a `.tar.gz` or similar input must be decompressed before it reaches Casita. The root is published only after the complete archive has been validated and converted to a canonical directory graph.
```rust
use casita::{Repository, RootName, import::TarImport, TarImportLimits};
use tokio::io::BufReader;
let repository = Repository::local("./cache").await?;
let archive = tokio::fs::File::open("release.tar").await?;
let report = repository
.import(
TarImport::new(BufReader::new(archive), RootName::try_from("releases/current")?)
.with_limits(TarImportLimits::default()),
)
.await?;
println!("{} entries in {}", report.entries, report.root);
```
`TarImportLimits` bounds archive bytes, entries, one file, total file bytes, sparse expansion, and in-flight files (`max_in_flight_files`, default 16). Set lower deployment-specific limits before accepting untrusted archives.
# Cargo
> How the experimental Cargo fork stores dependency data and build artifacts in Casita.
The experimental Cargo integration adds an `ArtifactStorage` interface inside Cargo, with filesystem and Casita implementations. Casita remains a generic repository: Cargo sends local paths through IPC, and Casita imports or restores their content. It does not gain Cargo-specific lockfile, package-resolution, or checksum object formats.
> **Experimental fork**
>
> This integration is implemented on the [`artifacts+casita` branch of the Cachix Cargo fork](https://github.com/cachix/cargo/tree/artifacts%2Bcasita) and is not part of an upstream stable Cargo release. It requires the fork’s unstable `-Zcasita-storage` flag and the pre-release Casita CLI.
## Enable the backend
Install the Casita CLI, which includes local IPC support:
```console
$ cargo install --path /path/to/casita/crates/casita
```
Select the backend in Cargo configuration:
**.cargo/config.toml**
```toml
[cache]
storage = "casita"
```
Then run the patched nightly Cargo with the unstable feature enabled:
```console
$ cargo build -Zcasita-storage
```
Cargo uses the default per-user Casita repository, normally the operating system’s data directory joined with `casita`. If its local service is not already listening, Cargo starts:
```console
$ casita --repository /casita ipc
```
## What changed in Cargo
The fork makes these Cargo-side changes:
1. `ArtifactStorage` lets Cargo prepare and persist registry archives, indexes, extracted sources, Git databases and checkouts, and workspace build state. Its filesystem implementation preserves Cargo’s usual layout.
2. Cargo’s `[cache] storage` selector accepts `"casita"`, gated by the unstable `-Zcasita-storage` feature.
3. `CasitaArtifactStorage` implements the same interface through bounded local JSON-RPC calls to `casita ipc`. Cargo does not link Casita’s async repository implementation.
4. The client negotiates protocol version 1 and requires the `artifact.checkout` and `artifact.import` capabilities before moving data.
5. Cargo uses writable local checkouts, asks Casita to restore retained content before use, and imports updates after Cargo changes them.
This keeps ownership clean:
```text
Cargo resolution, checksums, extraction, and builds
↓
Cargo artifact-storage adapter
↓
local artifact.checkout / artifact.import
↓
generic Casita roots and storage
```
## Stored artifacts
| Cargo data | Behavior |
| --------------------------------------------- | ---------------------------------------------------------------------- |
| Downloaded registry `.crate` archives | Imported individually and restored on demand |
| Registry indexes and extracted sources | Retained under names chosen by the Cargo adapter |
| Git databases and checkouts | Retained under names chosen by the Cargo adapter |
| Workspace target and intermediate build state | Restored into writable checkouts and retained per workspace and layout |
The adapter chooses root names and can use Casita’s filesystem, tar, and native Git importers for the corresponding content. Cargo still owns the meaning of each archive, checkout, and build output.
Explicit `--target-dir`, `CARGO_TARGET_DIR`, `build.target-dir`, and `build.build-dir` settings continue to take precedence. When one of those settings applies, Cargo preserves the explicitly selected filesystem location instead of silently redirecting it into Casita.
## Lifecycle and recovery
A later Cargo process can restore retained dependency data and workspace artifacts from Casita. The integration test deletes Cargo’s unpacked registry sources and confirms that an offline build recovers the `.crate` archive from Casita before extracting it again.
`cargo clean -Zcasita-storage` operates on the private workspace checkout and imports the now-empty directory. A later build therefore recompiles instead of recovering artifacts that were present before the clean.
Casita-managed roots do not participate in Cargo’s global-cache garbage collector. They remain subject to Casita’s own root and collection model. Until the integration exposes retention controls, operators should treat that as a separate lifecycle and storage-budget decision.
## Current limits
* The integration is a fork experiment, not an upstream Cargo compatibility promise.
* Its import and restore path is not yet on par with Cargo’s filesystem backend for performance. Build speed is a requirement before recommending it for regular use.
* It stores dependency data and mutable workspace build state; complete lockfile/registry evidence bundles are outside this integration.
* It does not reuse compiled artifacts across workspaces.
* IPC is local to the current machine and user. It is not a remote build-cache protocol.
* Cargo still owns package resolution, checksum validation, archive extraction, freshness decisions, compilation, and diagnostics.
* Casita verifies and retains the resulting filesystem graphs; it does not decide whether a Cargo artifact is semantically valid or reusable.
Read [Local IPC](../ipc/) for the framing, capability, and local trust boundary used by the adapter.
# Local IPC
> Local JSON-RPC protocol for importing and restoring artifacts.
The `casita ipc` service lets a local client import or restore artifacts by name. Requests contain paths on the daemon’s machine. The daemon reads or writes the files; artifact bytes do not pass through JSON.
| Method | Purpose |
| -------------------------------------------- | ------------------------------------------------------------------------- |
| [`rpc.initialize`](#initialize) | Negotiate protocol version and frame size. Call first on each connection. |
| [`artifact.import`](#import) | Store an artifact under a durable root. |
| [`artifact.restore`](#restore) | Restore a root to a local path. |
| [`artifact.checkout`](#checkout) | Check out a directory root, including a cache-miss response. |
| [`rpc.shutdown`](#shutdown-and-cancellation) | Close this connection. |
## Run the service
```console
$ cargo install --path crates/casita
$ casita --repository ./cache ipc
```
Without `--repository`, the service uses the CLI’s per-user repository. Only one service can listen for a repository at a time. Build with `--features git` to enable Git import and restore.
## Transport and framing
The endpoint is a Unix-domain socket or Windows named pipe. Unix socket and parent-directory permissions restrict access to the current user.
The endpoint suffix is the first 16 hexadecimal characters of the BLAKE3 hash of the repository path’s lossy UTF-8 representation. Unix uses `/tmp/casita/cargo-\.sock`; Windows uses `\\.\pipe\casita-cargo-v\`. Use the same repository path spelling as the daemon when computing the endpoint.
Send one JSON-RPC 2.0 object per UTF-8 line, ending with LF (`\n`). Responses use the same framing. CRLF, blank lines, and multiline JSON are invalid. The JSON examples below represent individual lines; send an LF after each one. Use absolute artifact paths to avoid dependence on the daemon’s working directory.
Requests on one connection run in order. Include an `id` to receive a response with the same `id`; failures use `error` instead of `result`. Separate connections can operate concurrently.
## Connection limits and deadlines
| Limit | Default | When exceeded |
| ----------------------------------------------------------------------------------------------- | ---------- | ----------------------------------------------------------------------------------------------- |
| Simultaneous connections, idle ones included | 64 | The new client receives one `-32000` error frame with a `null` `id`, then the connection closes |
| Time to deliver one complete request line, counted from the previous response (or from connect) | 60 seconds | The connection closes without a JSON-RPC response |
| Time for the client to drain one response | 30 seconds | The connection closes; an import that already returned its result has still been committed |
The receive deadline includes idle time. Reconnect and call `rpc.initialize` again after a connection closes. Artifact operations themselves have no deadline; the receive timer restarts after a response.
`casita ipc --max-connections`, `--frame-timeout-secs`, and `--response-timeout-secs` change these defaults; every value must be positive. The service is part of the CLI.
## Initialize
Call `rpc.initialize` first on every connection. `versions` must contain `1`. Optional `max_frame_bytes` defaults to 1,048,576 (1 MiB) and accepts 4,096 through 1,048,576 bytes, excluding LF. The server rejects longer requests; this limit does not currently bound responses.
```json
{"jsonrpc":"2.0","id":1,"method":"rpc.initialize","params":{"versions":[1],"max_frame_bytes":1048576}}
{"jsonrpc":"2.0","id":1,"result":{"version":1,"max_frame_bytes":1048576,"capabilities":["artifact.checkout","artifact.import","artifact.restore"],"importers":["filesystem","blob","copy","nar","filesystem_nar","tar","casitar"]}}
```
`importers` also contains `"git"` when the daemon has the `git` feature. Initialization succeeds only once per connection.
## Import
`artifact.import` takes named parameters. `importer` selects the input type (default: `filesystem`); `options` holds importer-specific settings. Required fields are listed below. Fields may be flat or grouped under `parameters`:
```json
{"jsonrpc":"2.0","id":2,"method":"artifact.import","params":{"importer":"blob","parameters":{"path":"/build/result.bin","root":"build/blob"},"options":{}}}
```
Keep `importer` and `options` outside `parameters`, and do not repeat a field in both places. Unknown importers, fields, and options fail before mutation. Every successful import publishes a durable root (Git uses `git/\`), so the result survives source deletion, restart, and collection. A failed import does not replace a root, but may leave collectible staged objects.
### Filesystem
Required fields: `path` (directory) and `root` (destination root name).
```json
{"jsonrpc":"2.0","id":2,"method":"artifact.import","params":{"importer":"filesystem","path":"/build/output","root":"build/latest","options":{"reread":true,"exclude":".control"}}}
{"jsonrpc":"2.0","id":2,"result":{"object":""}}
```
* `reread`: boolean, default `true`, preserving the original IPC behavior. Set `false` to allow reuse of unchanged files from the local ingest cache.
* `exclude`: optional single exact relative path to omit, not a glob or list.
The import publishes the directory under `root`. Legacy parameters such as `{"root":"build/latest","path":"/build/output"}` remain accepted, with the same result shape.
### Tar
Required fields: `path` (a tar or gzip-compressed tar file) and `root`. The daemon streams the archive from disk without extracting it first. `options.compression` is `"none"` (default) or `"gzip"`; filenames do not select compression. Gzip decompression streams directly into the tar importer. All gzip members and their checksums must validate before publication. Truncated members, corrupt trailers and trailing non-gzip bytes fail.
For gzip, `options.max_compressed_bytes` separately bounds compressed input (default 1 TiB). It is rejected with compression `"none"`. `options.limits.max_archive_bytes` always bounds decompressed tar bytes, including padding through EOF. Neither limit silently truncates the stream. For example:
```json
{"jsonrpc":"2.0","id":3,"method":"artifact.import","params":{"importer":"tar","path":"/archives/build.tar.gz","root":"build/from-gzip","options":{"compression":"gzip","max_compressed_bytes":104857600,"limits":{"max_archive_bytes":1073741824}}}}
```
```json
{"jsonrpc":"2.0","id":3,"method":"artifact.import","params":{"importer":"tar","path":"/archives/build.tar","root":"build/from-tar","options":{"limits":{"max_entries":10000}}}}
{"jsonrpc":"2.0","id":3,"result":{"object":"","archive_bytes":2048,"entries":1,"files":1,"directories":0,"symlinks":0,"hardlinks":0,"file_bytes":14,"sparse_expansion_bytes":0}}
```
`options.limits` accepts any subset of these nonnegative integer fields:
| Field | Default | Meaning |
| ---------------------------- | ---------------------- | -------------------------------------------- |
| `max_archive_bytes` | 1 TiB (1099511627776) | Complete raw tar bytes |
| `max_in_flight_files` | 16 | Concurrent file staging operations, positive |
| `max_entries` | 1,000,000 | Logical entries, excluding extension records |
| `max_path_bytes` | 4096 | Bytes in one archive pathname |
| `max_file_bytes` | 256 GiB (274877906944) | Logical bytes in one regular file |
| `max_total_file_bytes` | 1 TiB (1099511627776) | Total logical regular-file bytes |
| `max_sparse_expansion_bytes` | 256 GiB (274877906944) | Total materialized sparse holes |
The result includes the imported object and counts of consumed archive bytes, accepted entries by type, logical file bytes, and materialized sparse holes.
### Casitar
Required fields: `path` (a Casitar archive) and a nonempty `destinations` array. Supply one destination root name per archive root, in canonical archive-header order. The importer verifies the archive and publishes all destinations in one revision.
```json
{"jsonrpc":"2.0","id":4,"method":"artifact.import","params":{"importer":"casitar","path":"/archives/build.casitar","destinations":["build/restored"],"options":{"conflict_policy":"require_absent","limits":{"max_payload_bytes":1073741824}}}}
{"jsonrpc":"2.0","id":4,"result":{"mappings":[{"index":0,"root":"","name":"build/restored"}],"destination_revision":"","records_inserted":3,"records_reused":0,"payloads_written":1,"payloads_reused":0}}
```
`options.conflict_policy` is one of:
* `require_absent` (default): every destination must be absent.
* `replace_if_unchanged`: replace each destination only if its value remains unchanged between import preflight and final publication.
`options.limits` accepts any subset of these nonnegative integer fields:
| Field | Default | Meaning |
| ------------------------- | ---------------------------------- | ------------------------------------------ |
| `max_header_bytes` | 4 MiB (4194304) | Encoded header body bytes |
| `max_record_bytes` | 256 MiB (268435456) | Bytes in one encoded logical record |
| `max_payload_bytes` | u64 maximum (18446744073709551615) | Plaintext bytes in one payload |
| `max_total_payload_bytes` | u64 maximum | Total plaintext payload bytes |
| `max_archive_bytes` | u64 maximum | Complete archive bytes |
| `max_payloads` | 1,000,000 | Distinct payload frames |
| `max_records` | 1,000,000 | Distinct logical record frames |
| `read_buffer_bytes` | 64 KiB (65536) | Copy/hash scratch buffer; must be positive |
Frozen format ceilings still apply. Omit unchanged limits, particularly the 64-bit maxima, if the client’s JSON number implementation cannot represent them exactly. The result reports root mappings, publication revision, and counts of inserted/reused records and written/reused payloads.
### Git
Available with the `git` build feature. Required fields: `path` (local working tree or bare repository) and `view` (view name below `git/`). This imports native Git objects and publishes `git/\`.
```json
{"jsonrpc":"2.0","id":5,"method":"artifact.import","params":{"importer":"git","path":"/src/project","view":"upstream","options":{"refs":["refs/heads/main"],"max_cached_pack_bytes":0}}}
{"jsonrpc":"2.0","id":5,"result":{"view":"","objects":42,"revision":""}}
```
| Option | Default | Effect |
| ----------------------- | ----------------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| `refs` | Local branches and tags | Exact canonical refs to include, such as `refs/heads/main`. An empty list uses the default. |
| `revisions` | `[]` | Full hexadecimal object IDs to pin under `refs/casita/pins/`. Abbreviations and revision expressions are invalid. |
| `max_cached_pack_bytes` | 8 GiB (8589934592) | Maximum exact source pack retained for full clones; `0` disables it. |
| `concurrency` | 16 | Positive number of concurrently staged objects. |
| `max_buffered_bytes` | 64 MiB (67108864) | Positive budget for decoded bytes held by staging futures. Larger objects run alone. |
Pinned revisions appear in restored Git repositories. The result reports the view key, distinct native objects visited, and publication revision.
### Blob
Required fields: `path` (file) and `root`. Options are empty. The importer streams and publishes the bytes. Result: `{"object":""}`. Restore with `importer: "blob"` to an absent file.
### Repository copy
`importer: "copy"` requires `path` (a local Casita repository), `source_root`, and `root` (the destination name). Options are empty. It copies the complete source graph before publishing the destination root and returns `{"object":""}`.
### NAR and filesystem-NAR
`importer: "nar"` requires `path` (one canonical NAR archive) and `root`. Options are empty. Malformed archives and bytes after the archive fail. `importer: "filesystem_nar"` requires `path` (directory, regular file, or symlink) and `root`. Its only option is `reread`, default `false`, matching the library importer. Both use the library NAR validation and measurement paths.
```json
{"jsonrpc":"2.0","id":5,"method":"artifact.import","params":{"importer":"nar","path":"/archives/result.nar","root":"nix/result"}}
{"jsonrpc":"2.0","id":5,"result":{"object":"","nar_size":120}}
```
The stored object is a directory envelope with one entry named `root`. This retains directories, files, and symlinks alike. `artifact.restore` with `nar` or `filesystem_nar` unwraps the original node; ordinary checkout exposes the envelope.
## Restore
`artifact.restore` takes `root`, destination `path`, and an optional `importer` (default `filesystem`). It reads retained content without the original source.
| Importer | Restored result |
| -------------------------------------- | ------------------------------------------------------------------------ |
| `filesystem`, `tar`, `casitar`, `copy` | Directory tree |
| `blob` | File |
| `nar`, `filesystem_nar` | Original directory, file, or symlink from the NAR envelope |
| `git` | Bare Git repository with selected refs and object format; requires `git` |
Select the result’s content kind even when the root came from `copy` or `casitar`. For Git, use root `git/\`.
```json
{"jsonrpc":"2.0","id":6,"method":"artifact.restore","params":{"importer":"git","root":"git/upstream","path":"/restore/upstream.git"}}
{"jsonrpc":"2.0","id":6,"result":{"present":true,"object":""}}
```
A missing root returns `{"present":false}` without creating a destination. The parent directory must exist. Directory results accept an absent path or an empty real directory. File and symlink results require an absent path. Existing nonempty destinations are rejected. Restoration stages into a sibling on the same filesystem and publishes only the completed result; failures clean up staged output and leave existing destination contents untouched.
## Checkout
`artifact.checkout` takes a directory `root` and destination `path`. It returns `present: true` after materialization. For a missing root, it creates an empty destination directory and returns `present: false`.
```json
{"jsonrpc":"2.0","id":6,"method":"artifact.checkout","params":{"root":"build/latest","path":"/build/restored"}}
{"jsonrpc":"2.0","id":6,"result":{"present":true}}
```
The parent must exist. For a present root, the destination may be absent or an empty real directory. For a missing root, it must be absent.
## Shutdown and cancellation
`rpc.shutdown` closes only this connection after returning `null`. Omit `params`; the daemon continues serving other clients.
```json
{"jsonrpc":"2.0","id":7,"method":"rpc.shutdown"}
{"jsonrpc":"2.0","id":7,"result":null}
```
The `rpc.cancel` notification is accepted but currently does nothing. It does not interrupt an import or checkout.
## Errors
```json
{"jsonrpc":"2.0","id":8,"error":{"code":-32602,"message":"unsupported compression","data":{"category":"unsupported_option","importer":"tar","option":"compression"}}}
```
| Code | Meaning |
| -------- | ----------------------------------------------------------------------------------------- |
| `-32700` | Invalid JSON |
| `-32600` | Invalid JSON-RPC request, missing initialization, or unsupported/duplicate initialization |
| `-32601` | Unknown method |
| `-32602` | Invalid parameters, including unknown importer/options or incorrect field types |
| `-32008` | Import, checkout, filesystem, or repository operation failed |
| `-32000` | The connection limit was reached; sent with a `null` `id` before the connection closes |
Import and restore errors include `data.category`, `data.importer` and `data.option`. Importer is the requested name, or `null` if it could not be parsed. Option is the offending option name (for example `limits.max_entry`), or `null` when no specific option applies.
| Category | Code | Meaning |
| ---------------------- | -------- | ---------------------------------------------------------------------------- |
| `unsupported_importer` | `-32602` | Unknown importer or unavailable build feature |
| `unsupported_option` | `-32602` | Unknown import option or unsupported option value, such as compression |
| `invalid_parameters` | `-32602` | Missing fields, invalid names, wrong types or invalid parameter combinations |
| `execution_failure` | `-32008` | Source I/O, integrity, limits, destination conflict or repository failure |
Unsupported import requests are validated before source access or repository mutation, so callers can safely select another implementation. There is no required discovery handshake beyond normal `rpc.initialize`. Existing methods retain their numeric codes. Treat messages as diagnostic text, not stable machine-readable codes. Invalid framing (including oversized, empty, CR-containing, or unterminated lines) and an expired receive or response deadline close the connection without a JSON-RPC response.
The service does not expose listing/deleting roots, individual blob access, binary streaming, or garbage collection. Clients must verify that paths and root names are appropriate for their own trust boundary.
The historical [Cargo prototype](../cargo/) used this generic protocol. Cargo chose its root names and artifact lifecycle; the IPC service does not contain package resolution, build, registry, or checksum semantics.
# Use Casita as a Rust Library
> Open a repository, import data, and use the supported Rust API.
The supported API starts with `casita::Repository`. Open a persistent local repository with `Repository::local(path).await`, or use `Repository::memory()` for temporary data. The handle has no backend type parameters.
## Import and restore a directory
```rust
use casita::{Repository, RootName};
#[tokio::main(flavor = "current_thread")]
async fn main() -> Result<(), Box> {
let repository = Repository::local("./cache").await?;
let name = RootName::try_from("projects/demo")?;
let key = repository
.import(casita::import::FilesystemImport::new("./project", name))
.await?;
repository.checkout(&key, "./restored").await?;
repository.flush().await?;
Ok(())
}
```
`import` verifies the tree, stores it, and points the named root at its directory object. `checkout` needs an absent or empty destination. The Rust call does not create a checkout root. The import root above retains the graph.
Other built-in requests in `casita::import` handle blobs, tar, Casitar, Git (with `git`), and copies between repositories. See the [import guides](../guides/filesystem/) and [Rust API reference](../reference/rust-api/).
## Read a blob
```rust
use casita::{Repository, RootName};
use tokio::io::AsyncReadExt;
#[tokio::main(flavor = "current_thread")]
async fn main() -> Result<(), Box> {
let repository = Repository::local("./cache").await?;
let name = RootName::try_from("blobs/greeting")?;
let key = repository
.import(casita::import::BlobImport::new(&b"hello"[..], name))
.await?;
let mut reader = repository.open(&key).await?.ok_or("missing blob")?;
let mut bytes = Vec::new();
reader.read_to_end(&mut bytes).await?;
drop(reader);
repository.flush().await?;
assert_eq!(bytes, b"hello");
Ok(())
}
```
`open` returns a reader for a stored payload. It retains the data it needs against collection until dropped. Drop readers before a final `flush()` and await that flush before shutting down the runtime.
## Names, copies, and maintenance
| Task | API |
| ------------------------ | ------------------------------------------------- |
| Inspect a name or object | `root`, `roots`, `object` |
| Set or remove a name | `set_root`, `compare_and_set_root`, `remove_root` |
| Copy a named graph | `import(CopyImport::new(...))` |
| Export an archive | `export_casitar` |
| Check or collect | `fsck`, `preview_collection`, `collect`, `vacuum` |
`set_root` verifies the complete graph and replaces a name. `compare_and_set_root` replaces it only when its current value matches the expected value. A mismatch returns `false`. `remove_root` likewise requires the expected current key.
`CopyImport` reads one retained source root and copies its complete graph to the destination. The destination verifies it before setting its own root.
`fsck()` returns an integrity report. `is_healthy()` checks for reachable corruption; `is_clean()` requires no findings at all. Errors expose `kind()` and `retry_disposition()`. See [operations](../guides/operations/) and [errors](../reference/errors/) for details.
## Optional storage and advanced APIs
`Repository::s3(bucket, prefix, writer)` uses the optional, experimental `s3` storage profile. It uses standard AWS credentials. For a runnable upload and download example, see [`s3_sync.rs`](https://github.com/cachix/casita/blob/main/crates/casita/examples/s3_sync.rs).
Custom object formats, backends, transfer sources, and tuning APIs live under `casita::experimental` with the `experimental` feature. They may change between revisions. Start with [custom formats](../guides/custom-formats/) or the [experimental API reference](../reference/experimental-rust-api/). See [Cargo features](../reference/cargo-features/) for dependency settings.
## Test remote application workflows
These integration tests start a local S3-compatible server, so they need no AWS account or pre-existing bucket:
```console
$ devenv shell cargo test --features s3 --test s3_application_api
$ devenv shell cargo test --features s3 --test s3_multi_owner
$ devenv shell cargo test --features s3,experimental --test s3_multi_owner_recovery
```
They exercise uploads, downloads, retained readers, collection, and recovery across processes. The [multi-owner guide](../guides/s3-multi-owner/) explains the shared-repository model.
# What is Casita?
> Understand what Casita stores and when to use it.
Casita stores immutable objects and the links between them. Give an object graph a name, and Casita keeps everything reachable from that name. You can copy the graph to another repository, verify it there, and collect data that no name needs anymore.
It works with filesystem trees, native Git objects, IPLD blocks, and custom formats. Each format keeps its own identity. For example, a Git commit keeps its Git object ID. Casita provides the storage and lifecycle shared by those formats.
> **Pre-release**
>
> The CLI and Rust API may change. These docs follow the current `main` branch. Review [local repository recovery](../reference/local-repository/) before using Casita for important data.
## What Casita does
| Need | Casita’s role |
| --------------- | ------------------------------------------------------------------- |
| Store an object | Verify its bytes, identity, and links before publication. |
| Keep a graph | Point a named root at its first object. |
| Move a graph | Copy it to another repository, which verifies the received objects. |
| Reclaim space | Collect objects that no root or active operation retains. |
Casita stores the bytes separately from object identity. Chunking and compression can change without changing a Git ID, directory key, or other logical object key.
Casita does not decide which package to build, which Git branch to trust, or who may access a repository. The application using Casita makes those decisions. See [responsibilities](../concepts/responsibilities/) for the boundaries.
## Where to start
* [Quick Start](../getting-started/) stores and restores a directory.
* [CLI](../cli/) shows common commands.
* [Rust library](../library/) shows the supported application API.
* [Concepts](../concepts/) explains objects, roots, verification, and collection.
* [Reference](../reference/) covers exact syntax and advanced behavior.
The optional S3 storage profile and custom backends are experimental. See [Cargo features](../reference/cargo-features/) for build requirements.
# Reference
> Exact command, API, identity, format, feature, storage, and error contracts for Casita.
Use this section when you need the exact spelling or behavior of a Casita interface. For an introduction, start with the [Quick Start](../getting-started/), then read the task-oriented [CLI](../cli/) or [Library](../library/) guide.
## Reference map
| Topic | Use it to answer |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| [CLI](./cli/) | Which commands and flags exist, what they print, and which features they require |
| [Rust API](./rust-api/) | Supported application workflows, identities, errors, and reports |
| [Experimental Rust API](./experimental-rust-api/) | Custom backend composition, formats, sessions, and protocols |
| [Identifiers](./identifiers/) | How object keys, digests, roots, revisions, paths, refs, and SSH endpoints are validated |
| [Object formats](./object-formats/) | Which namespaces are built in, how they derive identity, and which links they retain |
| [Cargo features](./cargo-features/) | Build requirements for the CLI, native storage, Git, SSH, S3, and experimental APIs |
| [Local repository](./local-repository/) | What the standard persistent profile stores and how processes coordinate |
| [Benchmarks](./benchmarks/) | How end-to-end performance results are generated, validated, and compared |
| [Errors and integrity](./errors/) | How to classify failures, decide about retries, and interpret `fsck` |
## Contract levels
Casita documents three different kinds of fact:
* **Frozen encodings and semantics** are compatibility contracts. The [object formats reference](./object-formats/) summarizes them, and their golden vectors in the test suite pin the exact byte layouts.
* **Public Rust and CLI interfaces** describe the current pre-release revision. They may evolve even when the frozen data they read remains compatible.
* **Physical implementation details**, such as chunk sizes, compression, and local filenames, are operationally useful but are not logical object identity. They are labeled as private where they appear.
When a summary here and a normative specification disagree, use the specification for durable encoding questions and report the documentation discrepancy.
## Exact API documentation
The crate-level API map and item documentation live with the Rust source. Build the complete local Rustdoc, including feature-gated items, with:
```console
$ cargo doc --all-features --no-deps --open
```
The [Rust API reference](./rust-api/) explains how those items compose; it does not duplicate every method signature.
# Benchmarks
> Run and interpret Casita's reproducible performance suites.
The [benchmark harness](https://github.com/cachix/casita/tree/main/benchmarks) measures import, checkout, sync, verification, and collection. It also has focused suites for repository scale, Git history, catalog indexing, cache pressure, and network conditions. The [benchmark dashboard](/benchmarks/) shows published results.
Comparisons with Git, restic, Borg, and tar+zstd cover only operations the tools share. Their verification, retention, and storage semantics differ, so read the notes beside each result before comparing timings.
## Reproduce a run
From the pinned development environment, start with a short smoke run:
```console
$ devenv shell
$ benchmark --profile smoke \
--implementations casita,git,tar-zstd \
--cache-policies warm \
--repetitions 1
```
Each timed sample gets a fresh repository. Setup and correctness checks happen outside the timing window. The runner checks restored file bytes, paths, executable bits, and symlink targets before accepting a sample.
For publication, use the [full command and workload matrix in the benchmark README](https://github.com/cachix/casita/blob/main/benchmarks/README.md#publication-run). It uses both cache policies, all comparators, at least ten repetitions, `--require-all`, and `--require-clean`. Raw JSON records the source revision, tool versions, environment, inputs, and individual samples. Keep that JSON with its generated Markdown report in `benchmarks/baselines/`.
To regenerate a report without rerunning a benchmark, use `benchmark --render-existing RESULT.json`. Exploratory output belongs in `benchmarks/results/`; `benchmark-dashboard` builds the public page from the selected results.
## Graph traversal and spill
`graph-traversal` measures verification and collection when traversal state spills to temporary SQLite databases:
```console
$ benchmark run graph-traversal --profile smoke --repetitions 1 \
--output benchmarks/results/graph-traversal-smoke.json
```
The run records time, peak memory, spill size, and cleanup. Its correctness gate requires a forced spill and checks that temporary files are removed. See the [suite details](https://github.com/cachix/casita/blob/main/benchmarks/README.md#graph-traversal-and-spill) for the tested sizes and limits.
## Native Git scale
`git-scale` measures import and Git service operations across histories with many objects, deltas, large packs, and wide trees:
```console
$ benchmark run git-scale --profile smoke --shape many-objects
```
The optional huge profiles require a dedicated volume. Synthetic cases isolate individual costs; a real large repository provides a separate check. See the [Git scale suite](https://github.com/cachix/casita/blob/main/benchmarks/README.md#native-git-scale-suite).
## Catalog scale and S3 requests
`benchmark run catalog-index` measures catalog encoding, lookup, publication, reopening, and rebasing. It records memory, transferred bytes, and object-store requests. Its scale projections include request costs as well as catalog size. For current thresholds and assumptions, use the [catalog suite documentation](https://github.com/cachix/casita/blob/main/benchmarks/README.md#persistent-pack-index-benchmark).
## Retained history, cache pressure, and network limits
These suites test different sources of slowdown:
| Suite | What it varies |
| ------------------ | ---------------------------------------------------------- |
| `history-scale` | Retained generations and small updates |
| `pack-cache-scale` | Working set size and read pattern against a bounded cache |
| `network-scale` | Latency and per-connection bandwidth for S3 and atomic RPC |
Run a bounded check with `benchmark run SUITE --profile smoke --repetitions 1`. The [suite documentation](https://github.com/cachix/casita/blob/main/benchmarks/README.md#history-cache-and-network-scale-matrix) explains the workloads, correctness checks, and limits of local network and cache measurements.
# Cargo Features
> Compile-time feature flags, implications, and supported Casita capability sets.
Casita separates portable data-model and format code from the native repository profile and optional frontends.
## Feature table
| Feature | Implies | Adds |
| -------------- | ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------- |
| `native` | — | Tokio-backed repository workflows, persistent and memory backends, collection, transfer, filesystem I/O, and native Git view operations |
| `experimental` | — | Public `casita::experimental` namespace for backend, format, tuning, and protocol APIs |
| `cli` | `native`, `experimental` | The `casita` command-line binary, including the local JSON-RPC/NDJSON service (`casita ipc`) |
| `git` | `native` | Import from a local working tree or bare Git repository through `gix` |
| `oci` | `native` | Stream registry images into OCI image layouts and optional merged filesystems through `oci-client` |
| `git-fetch` | `native` | Experimental read-only Git fetch planning and pack generation |
| `git-http` | `git`, `git-fetch` | Read-only Git smart-HTTP serving and Tokio networking |
| `ssh` | `native` | Authenticated transfer sources through the system OpenSSH client |
| `s3` | `native` | Experimental shared S3 storage profile and `Repository::s3` |
| `fuzzing` | `native`, `experimental` | Internal parser adapters for fuzz harnesses |
The default feature set is `cli`, which includes `native` and `experimental`. An ordinary `cargo build` builds the `casita` binary. Library consumers that do not need the CLI can use `default-features = false, features = ["native"]`.
Library consumers using `default-features = false` also enable `experimental` to access service APIs such as Git fetch, HTTP, and SSH. Optional features alone do not expose backend or protocol types at the crate root.
## Common configurations
Casita has not published a tagged release. The examples below follow current development; replace `branch = "main"` with `rev = "..."` for reproducible evaluation.
```toml
# Standard library use with local repositories, without the CLI dependencies.
casita = {
git = "https://github.com/cachix/casita",
branch = "main",
default-features = false,
features = ["native"],
}
# Portable identities and filesystem data types without native I/O.
casita = {
git = "https://github.com/cachix/casita",
branch = "main",
default-features = false,
}
# Local native Git import.
casita = {
git = "https://github.com/cachix/casita",
branch = "main",
default-features = false,
features = ["git"],
}
# OCI registry image import.
casita = {
git = "https://github.com/cachix/casita",
branch = "main",
default-features = false,
features = ["oci"],
}
# Shared S3 storage through the application API; the storage profile is experimental.
casita = {
git = "https://github.com/cachix/casita",
branch = "main",
default-features = false,
features = ["s3"],
}
# Read-only native Git serving; `git` is implied.
casita = {
git = "https://github.com/cachix/casita",
branch = "main",
default-features = false,
features = ["git-http", "experimental"],
}
# Fetch planning and pack generation without the HTTP adapter.
casita = {
git = "https://github.com/cachix/casita",
branch = "main",
default-features = false,
features = ["git-fetch", "experimental"],
}
```
`Repository::s3` does not require the `experimental` feature. Custom backend composition and the generic repository require `experimental`, including when using local storage.
Install CLI combinations from a source checkout with:
```console
$ cargo install --path crates/casita
$ cargo install --path crates/casita --features ssh
$ cargo install --path crates/casita --features git
$ cargo install --path crates/casita --features oci
$ cargo install --path crates/casita --features git-http
```
## What remains without `native`
With `default-features = false`, the crate still provides:
* digests, typed IDs, object keys, records, roots, and revisions;
* canonical filesystem directory and path data types, including encoding and decoding.
Adding `experimental` also exposes `ObjectFormat`, `FormatRegistry`, portable payload verification, IPLD formats, native Git object/view models, and Casitar framing. Implementation modules remain private in every configuration.
Persistent stores, repository orchestration, streaming Casitar adapters, filesystem workflows, transfer execution, and service adapters require `native`. Native operations emit `tracing` spans and events but never install a global subscriber, leaving filtering and collection to the embedding application. `cli` adds the compact/JSON subscriber used by `--log-filter`, `--log-format`, and `RUST_LOG`; its default filter is `off`.
## CLI feature behavior
The CLI command tree always shows all Git subcommands. Their runtime feature requirements differ:
| Operation | Required build features |
| --------------------------------------------------- | -------------------------- |
| General CLI, Git view inspection, Git tree checkout | `cli` |
| `import -i git` | `cli,git` |
| `git serve` | `cli,git-http` |
| Local-to-local `sync` | `cli` |
| SSH-source `sync` and the remote source process | `cli,ssh` on both machines |
Calling `import -i git` or `git serve` from a binary without its required feature returns an explicit error rather than hiding the command from help output.
## Toolchain and generated documentation
The current minimum supported Rust version is 1.94.1 and the crate uses Rust edition 2024. `docs.rs` metadata enables every feature so feature-gated items appear together. Reproduce that API surface locally with:
```console
$ cargo doc --all-features --no-deps
```
# CLI Reference
> Command syntax, important defaults, and repository effects.
Install from a source checkout with `cargo install --path crates/casita`. Optional commands need the [corresponding Cargo feature](../cargo-features/). For a first walkthrough, start with the [CLI guide](../../cli/).
| Task | Commands |
| ------------------- | --------------------------------------------------------------------------------------------------------------------- |
| Store and read | [`import`](#import), [`object show`](#object-show), [`tree list`](#tree-list), [`cat`](#cat), [`checkout`](#checkout) |
| Retain and maintain | [`root`](#named-roots), [`gc`](#gc), [`vacuum`](#vacuum), [`fsck`](#fsck), [`holds`](#holds) |
| Move data | [`sync`](#synchronization), [`archive`](#portable-casitar-archives) |
| Integrate | [`run`](#run), [`ipc`](#ipc), [`git`](#native-git) |
## Global syntax
```text
casita [GLOBAL OPTIONS] COMMAND [COMMAND OPTIONS]
```
Put global options before the command. `-h`/`--help` prints command-specific help; `-V`/`--version` prints the binary version.
| Option | Effect |
| ------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| `--repository PATH` | Use this local repository instead of the per-user default. `sync` uses `--from` and `--to` instead. |
| `--log-filter DIRECTIVES` | Select trace levels; overrides `RUST_LOG`. Defaults to `RUST_LOG`, then `off`. |
| `--log-format compact\|json` | Write compact or newline-delimited JSON trace events to stderr. Default: `compact`. |
| `--spill-memory-objects OBJECTS`, `--spill-bytes BYTES` | Bound temporary traversal state. |
| `--pack-target-bytes BYTES` | Set the approximate compressed size of an immutable chunk pack. |
| `--pack-cache-bytes BYTES` | Set the S3 compressed-chunk cache size; `0` disables it. |
The pack cache option is intended for S3 endpoints used by `sync` or remote verified `cat`. Spill options affect temporary traversal state, not object identity.
For example, `casita=info` records operation outcomes and `casita=debug` adds internal decisions and contention. Casita’s own trace fields omit paths, root names, credentials, and payload contents, but dependency traces may have different rules. Use a scoped filter when sharing logs:
```console
$ casita --log-filter casita=info --repository ./cache fsck --dry-run
$ RUST_LOG=casita=debug casita --log-format json \
--repository ./cache gc --dry-run
```
Without `--repository`, Casita uses the OS application data directory, falling back to `$HOME/.casita` and then `.casita-data` when needed. It also searches the current directory and its parents for a `.casita` workspace marker. The marker scopes root names by a workspace UUID; bytes stay in the per-user repository. An explicit repository skips workspace scoping. The [`run` command](#run) uses the marker only for optional name shortcuts.
## Repository and filesystem commands
### `init`
```text
casita [--repository PATH] init
```
Without `--repository`, creates or reuses the current directory’s `.casita` workspace marker after opening the global profile. With `--repository`, creates or opens that local profile and prints its path and current repository revision.
### `import`
```text
casita [--repository PATH] import [-i IMPORTER] PATH [--root NAME] \
[--retention permanent|evictable] [--filesystem-rehash]
```
Without `-i`, Casita recognizes Git repositories from their metadata, probes regular-file headers for Casitar or tar, and otherwise imports a directory as a filesystem tree. Use `-i filesystem|tar|git|casitar|oci` to select an importer explicitly. Standard input (`-`) requires `-i`.
For a filesystem import, `--root` names the tree. When omitted, Casita derives a name below `auto/` from the canonical source path. In a workspace, both explicit and automatic names are scoped by its UUID. Importing the workspace directory omits its `.casita` marker. `--retention` publishes the root and policy together for filesystem and tar imports. Roots are permanent by default; omitting the flag keeps an existing root’s policy. Git, Casitar, and OCI imports do not accept this flag.
The filesystem importer prints the directory key and does not follow symlinks. It may reuse a file whose device, inode, size, and timestamps match the previous import. `--filesystem-rehash` reads every file again. Use `--filesystem-concurrency FILES` to change the number of files ingested at once (default 16), and `--chunk-upload-concurrency CHUNKS` to bound uploads per blob writer (default 32). See [Import semantics](../../concepts/imports/) for the reuse assumption.
### `import -i oci`
```text
casita [--repository PATH] import -i oci IMAGE --root NAME \
[--oci-platform OS/ARCH[/VARIANT]] [--oci-http] \
[--oci-max-blob-bytes BYTES] [--oci-max-total-blob-bytes BYTES] \
[--oci-rootfs-root NAME] [--oci-rootfs-max-bytes BYTES] \
[--oci-rootfs-max-entries COUNT]
```
Requires the `oci` Cargo feature. The importer selects one platform’s manifest and downloads its config and compressed layer blobs into a standard OCI image layout. Layer blobs stream into storage. It uses anonymous registry access and HTTPS by default.
`--root` names the OCI image layout, including its original layer archives. `--oci-rootfs-root` optionally names the merged container root filesystem, which can be checked out or mounted. The names must differ. Casita applies layer whiteouts and verifies uncompressed DiffIDs before publishing both roots atomically. The filesystem bounds limit decoded tar bytes and entries; the blob bounds limit the original compressed downloads. See [Import an OCI Image](../../guides/oci/) for the output and limits.
### `import -i tar`
```text
casita [--repository PATH] import [-i tar] FILE|- --root NAME \
[--retention permanent|evictable] \
[--tar-max-archive-bytes BYTES] [--tar-max-entries COUNT] \
[--tar-max-in-flight-files COUNT] \
[--tar-max-path-bytes BYTES] [--tar-max-file-bytes BYTES] \
[--tar-max-total-file-bytes BYTES] [--tar-max-sparse-expansion-bytes BYTES]
```
Streams one already-decompressed POSIX tar archive into the canonical filesystem model without extracting it to disk. A tar file is detected from its first valid header; use `-i tar` for standard input. The command accepts regular files, directories, symlinks, hard links, and old-GNU sparse files; PAX GNU sparse members are rejected. `FILE` may be `-` for standard input.
The root name is required and is committed only after complete validation. Defaults bound raw archive bytes (1 TiB), entries (1,000,000), path bytes (4096), individual files (256 GiB), total file bytes (1 TiB), and sparse-hole expansion (256 GiB). The corresponding flags override those bounds.
### `object show`
```text
casita [--repository PATH] object show KEY
```
`KEY` must be a full generic object key. Output includes the logical key, physical payload ID, payload size, ordered forward links, and current closure status.
### `tree list`
```text
casita [--repository PATH] tree list KEY
```
Lists the direct entries of a canonical filesystem directory. `KEY` may be a full `casita.directory.v1` key or a short `blake3-...` directory digest. The command requires a complete valid closure.
Entry output identifies directories (`d`), regular files (`f`), executable files (`x`), and symlinks (`l`).
### `run`
```text
casita [--repository PATH] run ROOT [--bin NAME_OR_PATH] [-- ARGS...]
```
Runs an executable from an existing named directory root. Casita searches the tree recursively and starts its sole executable. With several candidates, use `--bin` to choose a unique filename (`uv`) or relative path (`release/uv`); `./uv` selects a root-level file. It fails when no executable is found.
On Unix, candidates need an executable mode bit. On Windows, they need a `.exe` or `.com` extension. Internal executable symlinks can be selected, but discovery skips broken links, links outside the tree, and directory symlinks. Selection does not search `$PATH` or prefer a build profile.
```console
$ casita run cargo/builds/uv -- --version
$ casita run go/builds/server --bin server -- --port 8080
```
To shorten names, add run settings to the `.casita` marker created by `casita init`, keeping its header and workspace UUID:
```text
casita-workspace-v1
workspace = "12345678-1234-4234-8234-123456789abc"
[run]
default-scope = "cargo"
[run.scopes]
cargo = "cargo/builds"
go = "go/builds"
```
| Input | Resolved repository root |
| ----------------- | --------------------------------------------------- |
| `cargo:uv` | `cargo/builds/uv` |
| `go:server` | `go/builds/server` |
| `uv` | `cargo/builds/uv`, using the optional default scope |
| `cargo/builds/uv` | `cargo/builds/uv`, unaffected by the default |
| `/uv` | Literal root `uv`, bypassing all shortcuts |
Only bare names use the optional default scope. `SCOPE:NAME` selects a configured prefix, and a leading `/` requests a literal name. The closest marker wins; invalid scopes or configuration fail instead of falling back to another name.
`run` resolves the root once and retains that exact graph while the application runs. It materializes the tree under the repository’s `runs/` directory and normally removes it on exit. A crash may leave temporary files there. The child inherits the caller’s working directory, environment, and standard streams. Arguments after `--` go to the child without a shell. Casita returns its exit code and forwards common termination signals on Unix. Execution uses the caller’s permissions. See the [run guide](../../guides/run/) for a complete workflow and binary-selection details.
### `ipc`
```text
casita [--repository PATH] ipc [--max-connections COUNT] \
[--frame-timeout-secs SECONDS] [--response-timeout-secs SECONDS]
```
Starts the local JSON-RPC service for artifact import and restore. Defaults are 64 connections, 60 seconds to receive a request, and 30 seconds to write a response. See [Local IPC](../../integrations/ipc/) for the endpoint and protocol.
### `cat`
```text
casita [--repository PATH] cat KEY [--verified [--from ENDPOINT]]
```
Writes a blob’s exact plaintext payload to standard output. `KEY` may be a full generic key or a short `blake3-...` blob digest. Diagnostics go to standard error, so stdout may be redirected safely. `--verified` authenticates blocks before writing them. With `--from`, it reads a verified raw blob from a local, SSH, or S3 repository instead of the selected local repository; `--from` requires `--verified` and a raw blob key.
### `checkout`
```text
casita [--repository PATH] checkout KEY DIR [--no-root]
```
Materializes a complete canonical directory into `DIR`. The target is created if absent and must otherwise be empty. Checkout first writes a sibling staging directory on `DIR`’s filesystem and renames it into place. A target that cannot represent exact stored names, for example because of case folding or Unicode normalization, fails without a partial checkout. `KEY` accepts a full directory key or a short directory digest.
Every directory, file, and link is created relative to one open handle on the staging root, so a component swapped for a symlink, junction, or other reparse point while the checkout runs cannot redirect a write outside it. `DIR` itself must be a real directory: a path that is already a link is refused rather than written through. The path leading to `DIR` is the caller’s own authority, and a link stored inside the materialized tree is created as a link, not followed. Windows refuses to materialize a stored link whose target is absolute.
By default, successful checkout also registers an `auto/checkout/...` root derived from the canonical destination path. That root retains the materialized closure until explicitly removed. `--no-root` skips this step.
## Named roots
Roots retain the complete forward closure of one exact object.
### `root set`
```text
casita [--repository PATH] root set NAME TARGET [--retention permanent|evictable]
```
Atomically sets or replaces `NAME`. `TARGET` may be a full object key or a short filesystem digest. A short digest is rejected as ambiguous if matching blob and directory records both exist. The target closure must be complete and valid. Roots are permanent by default. `--retention` sets the policy in the same commit as the root. Without it, replacing a root keeps its existing policy. Evictable retention requires the local metadata backend.
### `root retention`
```text
casita [--repository PATH] root retention NAME permanent|evictable
```
Changes the policy of an existing local root. A permanent root remains until explicitly removed. An evictable root may be released by local disk-pressure collection, after which its name becomes a cache miss.
### `root rm`
```text
casita [--repository PATH] root rm NAME
casita [--repository PATH] root rm --prefix PREFIX
```
The first form removes one exact name. The second removes every name equal to or below `PREFIX` on root-name segment boundaries and prints each removed name. The two selectors are mutually exclusive. Removing a root makes data eligible for collection; it does not immediately delete objects.
### `root ls`
```text
casita [--repository PATH] root ls [PREFIX] [--long]
```
Lists targets and names. An optional positional `PREFIX` restricts output to equal or descendant root names. `--long` also displays `permanent` or `evictable` for each root.
## Collection and integrity
### `holds`
```sh
casita holds PATH [--json]
casita holds s3://BUCKET/PREFIX [--json]
```
Lists online data pins, released pin history, collector ownership, prune fences, and deletion claims without waiting for repository admission. Local inspection requires an existing repository; S3 additionally requires the `s3` feature. `--json` emits an object with `collectors`, `state`, and `coordination` fields. Each ledger has its own revision and exact tokens. Pin scopes include snapshot generations or closure roots; catalogs are identified by digest and encoded size. S3 collector entries include diagnostic writer names and legacy exclusivity flags. Listing never releases an existing hold or declares its owner dead. See [S3 maintenance](/guides/s3-maintenance/) for diagnosis and explicit recovery.
### `gc`
```text
casita [--repository PATH] gc [--dry-run]
```
Collection starts from every named root and active data pin. `--dry-run` takes exclusive ownership, runs the same mark plan, and reports removable logical records, payloads, and chunks without changing state. Without it, Casita first commits the logical prune and then removes unreferenced physical data.
Collection can proceed during mutations, transfers, and retained reads, preserving the data protected by their pins. Collectors serialize with other collectors. Starting a mutation on the standard local profile attempts nonblocking collection when disk usage has reached 80%. After collecting existing garbage, it releases least recently used evictable roots and vacuums until disk usage falls below 75% or no eligible roots remain. Explicit `gc` and `vacuum` preserve all named roots regardless of their retention policy.
### `vacuum`
```text
casita [--repository PATH] vacuum
```
Runs collection and forces reclamation of garbage deferred inside sparse packs. Use this when ordinary collection has left physical pack space to reclaim.
### `fsck`
```text
casita [--repository PATH] fsck [--audit-only | --dry-run] [--source REPOSITORY]
```
`--audit-only` checks logical records, root closures, payload identity and size, format relations, manifests, chunks, and unreferenced residue without running the separate repair pass. It is the appropriate mode for routine integrity measurement and is mutually exclusive with `--dry-run` and `--source`.
First, `fsck` verifies every physical payload representation named by the current logical snapshot. It rebuilds a corrupt existing Bao outboard and, when `--source` names a local replica with the expected fully verified payload, repairs a missing or corrupt local representation. It then checks logical records, root closures, payload identity and size, format relations, manifests, chunks, and unreferenced residue. `--dry-run` reports physical repairs without writing.
The repair pass never changes object records or roots, never repairs an unrelated I/O failure, and never invents logical state from a payload scan. It pins source and destination data for the operation, preserving replacement inputs while collection and ordinary readers continue. `fsck` prints a summary and each deterministic finding. Reachable corruption makes the command fail. Collectible residue and unavailable format verifiers are reported but do not, by themselves, make the repository unhealthy. See [Errors and Integrity](../errors/).
## Portable Casitar archives
All archive commands enforce finite defaults: 1 TiB for the complete stream, 256 GiB for one plaintext payload, 1 TiB for total plaintext payloads, and 1,000,000 each for payload and record frames. Override them with `--max-archive-bytes`, `--max-payload-bytes`, `--max-total-payload-bytes`, `--max-payloads`, and `--max-records`. The frozen 4 MiB header and 256 MiB encoded-record ceilings still apply.
Successful commands print the complete-file BLAKE3 digest and structural counts. `--json` selects the stable `casita.archive.v1` report schema.
### `archive create`
```text
casita [--repository PATH] archive create \
[--root NAME]... [--object KEY]... \
--output FILE|- [--force] [--json]
```
At least one selector is required; named roots and exact generic object keys may be mixed and repeated. Casita resolves names and verifies every selected closure through one source snapshot, deduplicates their union, then writes payloads and records in canonical order. Equal root sets and logical closures produce equal archive bytes regardless of selector order or traversal spill.
File output is staged, synced, and atomically published. It refuses an existing destination by default, including one created concurrently; `--force` selects atomic replacement. `--output -` writes only archive bytes to stdout and sends the human report to stderr. `--json` and `--force` are rejected for stdout output.
### `archive inspect`
```text
casita archive inspect FILE|- [--json]
```
Consumes and hashes the complete stream, requiring canonical framing, payload identities and lengths, section ordering, record encodings, the explicit end marker, and strict EOF. A successful result is labeled `structural`: it does not claim namespace reproduction, exact reachability, or complete root closures. This command does not open the configured repository. `-` reads archive bytes from stdin.
### `archive verify`
```text
casita archive verify FILE|- [--json]
```
Runs the full receiver import algorithm in an isolated temporary repository. Success means every payload hash, namespace-produced record, direct-link relation, exact record/payload membership rule, and declared closure verified. The temporary repository is discarded, and the configured durable repository is never opened or mutated. `-` reads from stdin.
### `archive import`
```text
casita [--repository PATH] archive import FILE|- \
(--root NAME... | --root-prefix PREFIX) [--replace] [--json]
```
Destination naming is always explicit. Repeat `--root NAME` exactly once per archive root in the canonical order printed by `inspect`, or use `--root-prefix PREFIX` to map them to `PREFIX/0`, `PREFIX/1`, and so on. There is no default prefix and no digest-derived or `auto/` naming.
The command streams every payload through BLAKE3, reuses exact physical payloads when possible, stages records through the registered namespace verifiers, checks the exact declared union closure, and publishes every mapped root together in one repository revision. By default all names must be absent. `--replace` captures their values before ingestion and replaces them only if every value remains unchanged at publication. Failure publishes no mapped root and may leave only receiver-verified unrooted residue for ordinary collection.
### `import -i casitar`
```text
casita [--repository PATH] import [-i casitar] FILE|- \
(--casitar-root NAME... | --casitar-root-prefix PREFIX) [--casitar-replace]
```
This is the common importer-command form of Casitar restoration. A file with the Casitar header is detected automatically; use `-i casitar` for standard input. It has the same all-or-nothing destination mapping and verification behavior as `archive import`, but every Casitar-specific option is namespaced: `--casitar-root`, `--casitar-root-prefix`, `--casitar-replace`, `--casitar-max-archive-bytes`, `--casitar-max-payload-bytes`, `--casitar-max-total-payload-bytes`, `--casitar-max-payloads`, and `--casitar-max-records`. It does not accept `--root`; use the Casitar mapping options instead. See [Import a Casitar Archive](../../guides/casitar/) for examples and Rust usage.
## Synchronization
```text
casita sync --from ENDPOINT --to ENDPOINT \
[--from-blobs ENDPOINT] [--writer NAME] \
[--object KEY]... [--root NAME]... \
[--path PATH [--destination-root NAME]] [--shallow] [--incremental]
```
`--from` supplies roots and object records. `--from-blobs`, when present, supplies payloads from a second Casita repository; it must contain the exact bytes required by those records. Sources accept local paths, S3 URLs (`s3://BUCKET/PREFIX`), or SSH URLs (`ssh://[user@]host[:port]/absolute/path`). Destinations accept local paths or S3 URLs. S3 needs the `s3` feature; SSH sources need `ssh` on both machines and a remote `casita` on the SSH user’s `PATH`. `--writer` sets the diagnostic WAL writer name for S3; otherwise `CASITA_WRITER` or a generated name is used.
Select at least one `--object` or `--root`; both are repeatable. Objects use full keys and copy their complete forward closure unless `--shallow` is set. Roots always copy complete closures and move under the same names only after receiver verification. `--incremental` reuses complete destination closures without auditing their descendants at the source. Intermediate object batches may remain after failure, but requested roots do not move. Sync does not remove destination data.
`--path PATH` requires exactly one filesystem root and cannot be combined with `--object` or `--shallow`. Only the selected file or directory closure is copied; verified path ancestors stay at the source. Without `--destination-root`, the copied closure is unrooted and collectible. Inline symlinks have no standalone key and cannot be rooted.
See [Synchronize Repositories](../../guides/sync/) for examples and failure behavior.
## Native Git
### `import -i git`
```text
casita [--repository PATH] import [-i git] SOURCE \
[--git-view VIEW] \
[--git-ref FULL_REF]... \
[--git-max-cached-pack-bytes BYTES] \
[--git-concurrency COUNT] [--git-max-buffered-bytes BYTES]
```
Requires `cli,git`. `SOURCE` is a local working tree or bare repository and is detected from its metadata. When auto-detected, an omitted `--git-view` uses the source directory basename. Each `--git-ref` is a full canonical name such as `refs/heads/main`. With no explicit refs, Casita selects local branches and tags. The command verifies native objects, publishes one immutable view, and atomically sets `git/VIEW`. When one verified source pack exactly matches that view, Casita retains packs up to 8 GiB by default as a rebuildable, byte-for-byte full-clone cache. This can substantially accelerate large clones but may nearly double stored bytes for incompressible repositories. Set `--git-max-cached-pack-bytes 0` to disable the cache or lower the limit to fit the storage budget. `--git-concurrency` bounds staged objects (default 16); `--git-max-buffered-bytes` bounds decoded bytes held by staging futures (default 64 MiB). An object larger than the budget runs alone.
### `git show`
```text
casita [--repository PATH] git show VIEW
```
Prints the selected view key, object format, optional default ref, and every direct or symbolic ref. It is available in the standard `cli` build.
### `git checkout`
```text
casita [--repository PATH] git checkout TREE DIR [--skip-gitlinks]
```
Safely materializes one exact type-qualified native Git tree into an empty directory, under the same handle-rooted containment as `checkout` above. Gitlinks fail by default; `--skip-gitlinks` materializes them as empty directories. This command is available in the standard `cli` build.
### `git serve`
```text
casita [--repository PATH] git serve VIEW \
[--listen ADDRESS] [--max-pack-bytes BYTES] \
[--pack-compression-level 0..9]
```
Requires `cli,git-http`. The default listen address is `127.0.0.1:9418`. The command binds one immutable view for read-only Git smart HTTP, prints its clone URL as `http://\/\.git`, and serves until stopped. Generated packs default to a 512 MiB limit and zlib compression level 6.
## Exit status and diagnostics
Successful commands return exit status `0`. Usage errors return `2`; runtime failures return `1`. There are no per-category exit codes.
Runtime failures are written to stderr as:
```text
error[]:
```
Usage errors use `error: \`. Programs should prefer the Rust API’s typed errors when they need richer classification or retry guidance rather than parsing CLI display text.
# Errors and Integrity
> Stable repository error categories, retry guidance, closure states, and fsck findings.
Casita keeps machine-readable classification separate from display text. Rust callers should match typed variants or use category helpers rather than parse `Display` messages. The [Library guide](/library/) covers built-in workflows.
## Repository error categories
`RepositoryError::category()` returns a non-exhaustive `RepositoryErrorCategory`. `as_str()` provides the stable spelling used by the CLI.
| Category | Stable string | Meaning |
| --------------------- | ----------------------- | ------------------------------------------------------------------------------------------------- |
| `Absent` | `absent` | A requested object, root, record, or payload is not present |
| `InvalidInput` | `invalid_input` | Caller input is malformed, foreign to this repository, or over a configured limit |
| `InvalidData` | `invalid_data` | Supplied or stored object bytes fail identity, canonical encoding, link, or relation verification |
| `ImmutableConflict` | `immutable_conflict` | One exact immutable key is already associated with a different record |
| `StaleRevision` | `stale_revision` | A compare-and-swap expected an obsolete repository revision |
| `DestinationConflict` | `destination_conflict` | A filesystem checkout destination is occupied or otherwise conflicts |
| `Busy` | `busy` | A nonblocking operation cannot acquire required ownership now |
| `Unsupported` | `unsupported` | A namespace, format, backend capability, or build feature is unavailable |
| `Corrupt` | `corrupt` | Committed state violates a repository invariant |
| `CollectedDuringRead` | `collected_during_read` | An unheld best-effort read raced collection of unrooted data |
| `Backend` | `backend` | I/O, storage, state-engine, or other operational infrastructure failed |
The enum is non-exhaustive. Include a fallback arm when matching it.
## Retry guidance
`RepositoryError::retry_disposition()` and `casita::experimental::Error::retry_disposition()` return a non-exhaustive `RetryDisposition`:
| Disposition | Caller interpretation |
| ---------------------- | ----------------------------------------------------------------------- |
| `Never` | Repeating the unchanged request cannot fix the reported condition |
| `Retry` | Retry may succeed, normally with bounded exponential backoff and jitter |
| `RetryAfter(duration)` | Wait at least the supplied duration, then retry with normal bounds |
| `Unknown` | The backend did not provide enough typed information to decide |
`Busy`, stale revisions, typed payload or state-backend transient failures, throttling, selected network I/O errors, and storage-full state may be retryable. Invalid identities, immutable conflicts, malformed input, and missing data normally are not.
A retry disposition does not make a non-idempotent application operation safe to repeat blindly. Observe the operation’s commit result, root expectation, or destination state before retrying work with external side effects.
## Closure status
`verify_closure()` returns one `ClosureStatus` for an exact snapshot:
| Status | Meaning |
| ---------------------------- | -------------------------------------------------------------------------------------- |
| `Complete { objects }` | Every reachable record and payload exists and all intrinsic format relations pass |
| `Missing { from, missing }` | The requested object itself or the first canonical reachable boundary has no record |
| `Invalid { object, reason }` | An object’s payload, identity, encoding, links, or direct relation failed verification |
| `Unsupported { object }` | The object’s namespace has no registered verifier |
`Complete` reports how many distinct objects the traversal visited rather than the set itself: a complete closure may be larger than the process verifying it, so the traversal spills to local storage instead of keeping the set in memory. Any count is meaningful only with the repository revision that produced it. A later state may add a previously missing object or use a different format registry.
Roots may be set only over `Complete` closures. Existing records can be unrooted or temporarily incomplete while an import or transfer is staging, but no successful root publication exposes such a graph.
## Integrity reports
`Repository::fsck()` inspects one logical snapshot protected by an online pin. Collection can reclaim unrelated data during the scan. Admission returns `Busy` if it conflicts with collection; retry after completion or recover an interrupted collector first. `FsckReport` records the inspected revision and counts of roots, objects, and unique payloads, followed by deterministic findings.
### Dispositions
| Disposition | Meaning | `is_healthy()` |
| ------------- | ------------------------------------------------------------------------ | -------------- |
| `Corrupt` | Reachable state violates an invariant | `false` |
| `Collectible` | Valid unrooted logical or unreferenced physical residue may be collected | unchanged |
| `Unchecked` | Exact validation could not run because a verifier is unavailable | unchanged |
`FsckReport::is_healthy()` means no reachable corruption was found. `FsckReport::is_clean()` is stricter: it requires no findings of any kind. A repository containing only collectible residue is healthy but not clean. A repository with an unchecked namespace can be reported healthy, but that does not prove the unchecked object’s format validity.
### Issue kinds
| Kind | Typical interpretation |
| ---------------------- | --------------------------------------------------------------------------------- |
| `StateEncoding` | Primary state could not be decoded or enumerated |
| `MissingRecord` | A root target or stored forward link has no logical record |
| `MissingPayload` | A logical record names physically absent payload bytes |
| `InvalidObject` | Payload identity, canonical encoding, recorded links, or a direct relation failed |
| `UnsupportedNamespace` | No verifier is registered for the namespace |
| `UnrootedObject` | A valid logical record is unreachable from every named root |
| `UnreferencedPayload` | A physical payload is referenced by no logical record |
| `UnreferencedChunk` | A physical chunk is referenced by no present payload |
The last three unrooted/unreferenced conditions are normally collectible, not reachable corruption.
## CLI behavior
The CLI prints runtime failures to stderr with the stable category supplied by the repository, Casitar, or frontend error:
```text
error[invalid_data]:
```
Usage failures use `error: \` and exit `2`. Runtime failures exit `1`; the stable category string is not a distinct numeric exit code. Success exits `0`.
`casita fsck` exits successfully when the report is healthy, even if it also reports `Collectible` or `Unchecked` findings. It fails when at least one `Corrupt` finding exists.
## Response guide
| Finding | First response |
| --------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| `busy` | Let the active mutation/read/collector finish, then retry with bounds |
| `stale_revision` | Read a fresh snapshot and recompute the conditional mutation |
| `collected_during_read` | Repeat under a `RetentionHold`, or root the data before relying on it |
| `unsupported` / `Unchecked` | Open with a registry or build that contains the required verifier |
| `Collectible` findings | Preview and run collection if the residue is no longer needed |
| `InvalidObject`, `MissingPayload`, or another `Corrupt` finding | Preserve the repository, stop treating it as authoritative, and restore or re-import from a trusted source |
| `backend` | Inspect the underlying I/O/storage error and available capacity before applying typed retry guidance |
`fsck` does not rewrite reachable records or synthesize missing payloads. It can rebuild derived physical state or replace a bad physical representation only from an independently verified replica; it never changes logical records. Collection removes unreachable residue and is not a substitute for restoring corrupted reachable data.
# Experimental Rust API
> Custom backends, formats, sessions, transfer sources, and limits.
Enable `experimental` to compose Casita’s repository engine from your own payload store (`PS`), metadata store (`SS`), and immutable `FormatRegistry`. Import these types from `casita::experimental`. This API may change between revisions. For built-in workflows, use the [supported Rust API](../rust-api/) and [library guide](../../library/).
| Need | Section |
| ---------------------------------------- | --------------------------------------------------------- |
| Open or compose a repository | [Constructors](#constructors) |
| Stage verified objects and publish roots | [Publication](#publication) |
| Retain data while reading | [Stable reads and retention](#stable-reads-and-retention) |
| Implement storage or format behavior | [Backend traits](#backend-traits) |
| Set bounds and classify failures | [Limits](#limits), [Errors](#errors) |
Import requests and `Importer` live in `casita::import` for both APIs. `MultiRootFilesystemImport` and `UnrootedFilesystemImport` are available there with `experimental`.
Native operations need a Tokio runtime with timers; custom runtime builders should call `enable_all()`. Drop sessions and readers, then await `flush_repository_leases()` before runtime shutdown to finish durable pin releases.
Build item-level Rustdoc with:
```console
$ cargo doc --all-features --no-deps --open
```
## Constructors
| Constructor | Result |
| ------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------ |
| `Repository::memory()` | Ephemeral `MemoryBlobStore` plus `MemoryMetadataStore` with built-in formats |
| `Repository::local(path).await` | Standard persistent `ChunkedBlobStore` plus `TursoMetadataStore`, including cross-process coordination |
| `Repository::new(payloads, state)` | Caller-supplied backends, built-in formats, and default deployment limits |
| `Repository::with_formats(payloads, state, formats, limits)` | Caller-supplied backends, exact immutable registry, and explicit `FormatLimits` |
`Repository::with_fs_coordination(path)` adds cross-process mutation/read versus collection ownership to a custom composition sharing one local root. Only use it when every process opening those backends agrees on that root.
### Deployment profiles
`RepositoryProfile` holds the deployment policy a repository applies on top of its backends. `Repository::with_profile(profile)` replaces a handle’s profile; existing clones keep theirs.
| Profile | Behaviour |
| -------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `RepositoryProfile::generic()` | What `Repository::new` starts from: no cross-process coordination, spills in the platform temporary directory, no import cache, no implicit maintenance |
| `RepositoryProfile::local(root).await` | What `Repository::local` applies: collector locks and spill files under `root`, emergency collection when storage is full, and a disk-pressure collection before a handle’s first mutation. Opening reserves the lock file and removes stale spill files |
`with_fs_coordination(root)`, `with_spill_limits(limits)` and `with_ingest_cache(&turso_metadata)` adjust a profile before it is applied. The ingest cache lets a re-import skip files whose identity is unchanged; it must belong to the same `TursoMetadataStore` the repository uses. The local profile assumes payloads and state share the filesystem holding `root`:
```rust
std::fs::create_dir_all(root.join("blobs"))?;
let payloads = ChunkedBlobStore::local_packed(root.join("blobs")).await?;
let metadata = TursoMetadataStore::open(root.join("casita.sqlite")).await?;
let profile = RepositoryProfile::local(&root).await?.with_ingest_cache(&metadata);
let repository = Repository::new(payloads, metadata).with_profile(profile);
```
`Repository::local` additionally takes the collector lock while it opens and migrates the payload catalog, and reclaims abandoned metadata before returning; prefer it when its fixed layout suits the deployment.
## Core logical types
Catalog construction and binary encoding are free functions in this namespace:
| Value | Construct | Encode / decode |
| -------------- | ---------------------------- | ----------------------------------------------- |
| `ObjectKey` | Public identity constructors | `encode_object_key` / `decode_object_key` |
| `RootRecord` | `new_root_record` | `encode_root_record` / `decode_root_record` |
| `ObjectRecord` | `new_object_record` | `encode_object_record` / `decode_object_record` |
`new_object_record` validates link ordering but does not verify the payload or graph. `LogicalEncodingError` and `ObjectRecordError` also live here. The frozen binary layouts are unchanged; enabling `experimental` does not add construction or encoding methods to the supported record types.
| Type | Meaning |
| ----------------------------------------------------- | ------------------------------------------------------------------------------------------- |
| `ObjectKey` | Versioned namespace plus canonical format-native identifier |
| `ObjectRecord` | Immutable key, BLAKE3 payload ID, payload length, and canonical forward links |
| `RootName` / `RootRecord` | Durable name selecting one exact object and retaining its complete closure |
| `RepositoryRevision` | Opaque token for one logical state; compare for equality only |
| `RepositoryGeneration` | Position of a logical state in its repository’s commit order; ordered within one repository |
| `ClosureStatus` | `Complete`, `Missing`, `Invalid`, or `Unsupported` result for one snapshot traversal |
| `Digest`, `BlobId`, `DirectoryId`, `ChunkId` | Raw and capability-typed BLAKE3 identities |
| `Directory`, `Node`, `PathComponent`, `SymlinkTarget` | Canonical filesystem data model |
See [Identifiers](../identifiers/) for their exact textual and validation rules.
## Publication
Call `Repository::mutation_session().await` to acquire a `MutationSession`. The standard local profile first attempts pressure-driven collection at 80% used capacity, then the session registers a staging pin before payload writes begin. Generic repository compositions skip the admission policy.
The main staging methods are:
* `stage_blob` and `stage_blob_reader` for raw payloads;
* `stage_object` for a caller-selected generic key and payload;
* `stage_directory` for canonical directory data;
* `stage_existing` when the payload already exists in the same store; and
* `stage_git_blob_file` and `stage_git_blob_files` to reuse stored native Git blobs as ordinary files without rereading their bytes (built-in formats only).
Staging returns `StagedObject`, a sealed value tied to that exact repository instance. Namespace verification has already reproduced its identity and links. An arbitrary `ObjectRecord` is not accepted as a publication substitute.
Publication choices include:
* `publish_unrooted` for verified records only;
* `publish_closures` for records plus bounded, checked closure targets without named roots;
* `publish_rooted` for records plus one root;
* `publish` for records plus a batch of `RootChange` values;
* `publish_at_revision` for an exact compare-and-swap; and
* `publish_if_roots_match` for linearizable compare-and-publish against exact current root values.
Records and root changes commit atomically. A root is published only after its resulting closure is complete and valid. Unrelated revision races can be retried; an observed root mismatch is returned as `ConditionalPublishResult::RootMismatch` without overwriting the changed name.
`publish_closures` verifies staged or existing targets with normal format and link checks before atomically publishing records and requested witnesses. Staged-object and target counts are each limited by `max_batch_objects`. Existing witnesses may be reused; this is a completeness check, not a fresh corruption audit. Only the targets gain witnesses, and a built-in raw or Git blob needs none because its record proves its own closure. The mutation retains the checked graphs for its lifetime, but their witnesses do not become permanent roots.
## Stable reads and retention
`Repository::retention_hold().await` returns a borrowed `RetentionHold` over one immutable `MetadataSnapshot`. `owned_retention_hold()` provides the owned variant for services that need a `'static` lifetime.
A hold provides snapshot object lookups, payload opens, and closure verification while preventing collection from removing physical data visible to that snapshot. Collection can reclaim unrelated data while the hold remains live. Dropping the hold makes its otherwise unreferenced data collectible.
Repository helpers such as filesystem checkout and transfer source sessions take the required hold internally.
## Repository workflows
| Workflow | Primary API | Extra capability |
| -------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------- |
| Exact-length object ingestion | `MutationSession::stage_object_reader_with_size` | `native`; verifies while writing, with no stored-payload reread |
| Raw blob import | `Repository::import(BlobImport::new(reader, root))` or `MutationSession::import(BlobImport::new(reader, root))` | `PS: BlobStore`, `SS: MetadataStore`, `native` |
| Filesystem import and checkout | `Repository::import(FilesystemImport::new(...))`, `FilesystemImport::new(...).reread(true)`, `Repository::checkout` | `PS: BlobStore`, `SS: MetadataStore`, `native` |
| Tar stream import | `Repository::import(TarImport::new(...))` | `PS: BlobStore`, `SS: MetadataStore`, `native` |
| Closure validation | `Repository::verify_closure` or a retention hold | `PS: BlobStore`, `SS: MetadataStore` |
| Logical collection | `preview_logical_collection`, `collect_logical`, `try_collect_logical` | `SS: MetadataStore` |
| Collection preview and execution | `preview_collection`, `collect`, `try_collect` | `PS: BlobGc` |
| Integrity inspection | `Repository::fsck` | `PS: BlobGc` |
| Physical fsck repair preflight | `Repository::fsck_repair`, `Repository::preview_fsck_repair` | `ChunkedBlobStore` |
| Generic transfer | `transfer` | Destination repository plus a `TransferSource` |
| Named graph import | `Repository::import(CopyImport::from_source(source, source_name, destination_name))` | Custom `TransferSource`; built-in sources use `CopyImport::new` |
| Path-selected transfer | `transfer_path` | Named filesystem root plus a `TransferSource` |
| Git view publish/read/checkout | `publish_git_view`, `read_git_view`, `checkout_git_tree` | `native` |
| Native local Git view import | `Repository::import(GitImport::new(...))` | `git` |
| Native Git closure import | `Repository::import(GitClosureImport::new(objects_dir, roots))` | `git`; retained result without a view |
| Session-scoped Git closure import | `GitClosureImport::new(objects_dir, roots).import(&session)` | `git`; receiving session retains the result |
| Git fetch service | `GitFetchService` | `git-fetch`; HTTP adapter requires `git-http` |
| SSH source | `SshTransferSource` | `ssh` |
| Casitar stream I/O, export, and import | `CasitarReader`, `CasitarWriter`, `Repository::export_casitar`, `Repository::import(CasitarImport::new(...))` | `native` |
## Backend traits
### Payload storage
`BlobStore` is the minimum physical interface: presence checks, seekable reads, streaming writes, and optional chunk metadata. Convenience methods provide whole-slice writes and reads.
Every implementation states its publication contract through the required `publication()` method. `PayloadPublication::Immediate` means each closed writer is durable. `PayloadPublication::Cataloged` points at a `CatalogPublication` implementation, which seals staged writes and coordinates the physical catalog with logical commits. An adapter forwards the capability of the store that receives its writes; there is no default to inherit.
`BlobGc: BlobStore` adds enumeration and deletion of payloads and chunks. It is required for physical collection and `fsck` because those operations compare logical reachability with physical coverage. Logical-only collection requires only `MetadataStore`; it prunes records without making a payload-liveness claim. Use it when a service shares physical blobs across repositories or tenants. Its `*_pinned` deletion methods have no defaults. Each store states how it honours online pins, and a wrapper forwards them together with any behavior it adds, so collection cannot silently bypass the pin ledger.
`BlobSync` is an optional destination capability exposed through `BlobStore::as_blob_sync`. A source needs only `BlobChunkSource` and its chunk map for transfer to negotiate missing chunks. Otherwise the same logical operation streams and re-verifies complete plaintext payloads.
Built-in implementations include `MemoryBlobStore`, `ChunkedBlobStore`, `CombinedBlobStore`, and `RepairingBlobStore`.
#### Near/far payload composition
`CombinedBlobStore::new(near, far)` is a public `BlobStore` implementation for tiered deployments. Reads and `has` checks prefer near and fall back to far; writes are opened only in near. A far read is not copied into near.
```rust
use casita::experimental::{BlobStore, CombinedBlobStore, MemoryBlobStore};
async fn read_from_far() -> Result<(), casita::experimental::Error> {
let near = MemoryBlobStore::new();
let far = MemoryBlobStore::new();
let existing = far.put_slice(b"upstream").await?;
let payloads = CombinedBlobStore::new(near.clone(), far);
assert_eq!(
payloads.read_to_vec(&existing).await?,
Some(b"upstream".to_vec()),
);
assert!(!near.has(&existing).await?); // no implicit cache warming
Ok(())
}
```
The far type still satisfies `BlobStore`, but the adapter never opens a writer on it. `CombinedBlobStore` intentionally implements neither `BlobGc` nor a combined `BlobSync`: it cannot establish global liveness for a shared far tier, and it falls back to verified whole-blob transfer. Applications that own those cross-tier policies can provide a custom backend implementing the additional capabilities.
`RepairingBlobStore::new(near, far)` is the opt-in corruption-aware composition for two `ChunkedBlobStore` values. It performs a complete validation before a reader can expose bytes. Typed near corruption or a missing referenced chunk causes one per-blob repair flight: the far payload is verified in full, its compressed chunks are independently checked while copied, and the near manifest is atomically replaced only after complete identity verification. Other backend failures are returned unchanged. `BlobRepairError` retains both near and repair-source diagnostics when recovery fails.
`RepairingBlobStore::verified_read` applies the same policy to Bao ranges and rebuilds derived outboards from verified near bytes. Its `BlobGc` capability owns only the near tier and excludes deletion while a local reader, writer, or repair is active; the far tier requires independent lifetime coordination.
### Repository metadata
`MetadataStore` has two foundational operations:
* `snapshot()` returns an immutable, internally consistent `MetadataSnapshot`;
* `commit(expected_revision, mutation)` atomically compares and applies one `MetadataMutation`.
`MetadataSnapshot` exposes its revision plus lazy object, root, and complete enumeration methods. A metadata backend is trusted infrastructure. Applications should normally mutate it through `Repository`, which verifies objects and root closures before constructing metadata mutations. `repository.metadata()` accesses the backend; `repository.payloads()` accesses the separate blob store.
Built-in implementations are `MemoryMetadataStore`, `TursoMetadataStore`, and, with `s3`, `Wal3MetadataStore`. Custom backends must provide a durable `PinStore`, `MetadataSnapshot::generation()` and `objects_created_through(generation)` for online retention. A successful commit advances the generation atomically; idempotent object inserts preserve their creation generation. Unsupported generation methods fail closed instead of retaining future garbage indefinitely. Lazy snapshots also report immutable files through `retention_resources()`.
`DataPinLease` protects scoped reads and staged writes while collection proceeds. `RepositoryLease` is collector-only ownership returned by `try_collection_lease()`. Both `try_collection_lease()` and `coordinates_payload_catalog()` are required: a store shared between processes without a common lock returns a durable lease, while local and in-memory stores return `RepositoryLease::process_local()`. Wrappers forward both, so they cannot silently drop a shared store’s collector exclusion or its payload catalog.
For an already-open repository, `recover_collection(&collector_token)` retries an interrupted physical collection using the exact token in its pin inventory. The caller must first establish that the collector and all requests covered by its claims have stopped. Recovery preserves live pins and keeps interrupted claims and prune fences through marking and commit. S3 operational ownership must be recovered separately before this call. A stale token cannot take over a newer collector. This API does not expire or release abandoned reader pins. For the standard S3 profile, `Repository::recover_s3_collection(bucket, prefix, writer, &collector_token)` also performs the reopen under collector ownership, including when an abandoned prune fence blocks ordinary read admission.
### Read-only chunk sources
`BlobChunkSource::get_chunk` reads one compressed chunk. It has no write, metadata, or collection operations. `BlobSync` extends it with destination presence checks and verified chunk/manifest writes.
`TransferReadSession::as_chunk_source()` exposes source reads; a destination continues to expose `BlobStore::as_blob_sync()`. A writable blob backend also gets `BlobStore::as_chunk_source()` by default. Transfer still copies only missing chunks and verifies their contents and complete payload identity before publishing records or moving roots. The session must protect the actual blob storage for the full operation; a read interface alone does not establish that retention agreement.
This prepares the API for independent blob transports. Session-bound S3 descriptors and direct-export retention are still proposed work.
Custom chunk backends implement `BlobChunkSource` separately from `BlobSync`. Import `BlobChunkSource` for chunk read method calls on concrete chunk stores.
### Object formats
Implement `ObjectFormat` to add one deterministic namespace. `verify` consumes a `VerificationContext`, checks the format’s native identity and canonical payload, and returns a sealed `VerifiedObject` with exact forward links.
Override `verify_links` when validity depends on an intrinsic relation to direct children, such as checksum evidence comparing a linked archive. The `DirectLinkView` exposes only the declared direct targets. Verification must be deterministic and must not depend on network or mutable policy.
Register `Arc\` values with `FormatRegistry::new`; duplicate namespace ownership is rejected.
### Transfer sources
`TransferSource::begin_transfer` opens a stable `TransferReadSession` bound to one revision and physical retention lifetime. The session exposes transfer- shaped record, root, payload, and optional chunk reads rather than leaking the source’s backend types. `path_proof` is an optional optimization: capable sessions return a retained-revision proof in one operation, while the default reports `PathProofResponse::Unsupported` and preserves the ordinary spine walk.
Both a local `Repository` and `SshTransferSource` implement this boundary. The destination never trusts sender-supplied links or identities: it reruns its own format verifier before publication.
`transfer_path` resolves one path beneath a named filesystem root. It verifies the directory spine without publishing those ancestors, transfers only the selected file or directory closure, and optionally installs that closure under an explicit destination-owned root. Both functions take `TransferOptions`, which selects the discovery policy. To transfer through an already-open stable session, for example after resolving roots on it, pass it as `&HeldSession(session)`: the held session keeps its revision and retention.
## Limits
`FormatLimits` configures hostile-input and deployment bounds. Defaults are:
| Field | Default |
| ------------------------------ | ---------- |
| `max_payload_bytes` | `u64::MAX` |
| `max_metadata_bytes` | 256 MiB |
| `max_links_per_object` | 1,000,000 |
| `max_directory_entries` | 1,000,000 |
| `max_batch_objects` | 4,096 |
| `max_root_changes` | 1,024 |
| `max_traversal_objects` | 10,000,000 |
| `read_buffer_bytes` | 64 KiB |
| `max_transfer_in_flight_bytes` | 64 MiB |
These are active repository limits, not durable identity parameters. Lowering a limit can make an otherwise valid large object unavailable to a deployment without changing that object’s key.
## Errors
Repository workflows return `RepositoryError`; physical and filesystem operations use `casita::experimental::Error`; transfer has `TransferError`. Do not classify failures from display strings. Use `RepositoryError::category()` and `retry_disposition()` where orchestration needs stable behavior. See [Errors and Integrity](../errors/).
# Identifiers
> Text forms and validation rules for Casita keys, digests, roots, revisions, paths, Git refs, and SSH endpoints.
Casita validates identifiers before they reach storage. Lengths below are byte lengths, not character counts.
## Digests and typed IDs
`Digest` is exactly 32 BLAKE3 bytes. Its canonical text form is:
```text
blake3-
```
The base64 alphabet uses `-` and `_`, never `+` or `/`, and carries no `=` padding. `BlobId`, `DirectoryId`, and `ChunkId` use the same text form but are distinct Rust types so callers must select the intended meaning.
`Digest::to_hex()` and `Digest::from_hex()` provide a lowercase 64-character hex form. Casita uses hex for sharded physical paths; the CLI’s short object form uses the `blake3-...` representation.
## Namespace IDs
A `NamespaceId` matches:
```text
label(.label)*.v
```
The complete value is ASCII and at most 64 bytes. Each ordinary label starts with a lowercase ASCII letter; the remaining bytes may be lowercase letters, digits, or `-`. The last component is `v` followed by one or more digits.
Valid examples include `casita.blob.v1`, `ipld.raw.v1`, and `git.sha256.commit.v1`.
## Object keys
An `ObjectKey` has the canonical text form:
```text
:
```
The namespace is at most 64 bytes and the decoded native ID is at most 128 bytes. The native ID is opaque to generic repository code: it is a BLAKE3 digest for several built-in formats, a CID for IPLD, and a native SHA-1 or SHA-256 OID for Git.
The `ObjectKey` parser does **not** accept `blake3-...`. Selected CLI commands accept that shorter spelling for filesystem blobs or directories and add the namespace themselves:
| Command position | Accepted form |
| -------------------------------------- | ------------------------------------------------------ |
| `object show KEY` | Full object key only |
| `sync --object KEY` | Full object key only |
| `tree list KEY` and `checkout KEY DIR` | Full key or short directory digest |
| `cat KEY` | Full key or short blob digest |
| `root set NAME TARGET` | Full key or an unambiguous short blob/directory digest |
If both a blob and directory record have the same short digest, `root set` requires the full key.
## Root names
A `RootName` is a durable UTF-8 name split into `/`-separated segments:
* the complete name is 1–1024 bytes;
* each segment is 1–255 bytes;
* leading, trailing, or repeated `/` is invalid;
* `.` and `..` segments are invalid; and
* ASCII C0 control bytes and DEL are invalid.
Other UTF-8 text, including spaces and non-ASCII characters, is permitted. Prefix operations use segment boundaries: `projects/demo` is under `projects`, but `projects-old` is not.
CLI filesystem imports without `--root` derive an `auto/...` name from the canonical source path. Checkout similarly derives `auto/checkout/...` unless `--no-root` is used. Treat these names as ordinary roots: they retain their complete closures until removed.
## Repository revisions
`RepositoryRevision` is 32 opaque bytes displayed as:
```text
rev-
```
A successful logical mutation creates a fresh revision. Revisions support equality checks only: they do not expose ordering and do not identify a state in another repository.
## Repository generations
`RepositoryGeneration` orders the states of one repository. Every successful commit advances it atomically with the new revision, and it never decreases, so of two readers of one repository, the one with the larger generation sees every commit the other sees. `MetadataReader::generation` and `RetainedReader::generation` report it; custom metadata backends without generations return an `Unsupported` error. It is displayed as:
```text
gen-
```
Generations of different repositories, or of a repository recreated at the same location, are unrelated.
## Filesystem names
`PathComponent` stores a raw byte name between 1 and 255 bytes. It rejects `/`, NUL, and the exact names `.` and `..`. Ordering is lexicographic over raw bytes and defines canonical directory order.
`SymlinkTarget` stores 1–4095 raw bytes. It rejects NUL but permits `/`, `.`, and `..`. This data-model validation does not imply that every target is safe to materialize; checkout applies its own platform and traversal protections. A stored target is link data rather than a path casita resolves: materializing it creates a link and never follows it, and containment comes from the open directory handle a checkout writes through.
## Git names
`CanonicalRefName` accepts full names below `refs/`, such as `refs/heads/main`. The frozen profile follows Git’s important safety rules: it rejects control and space bytes, forbidden punctuation and sequences, leading dot components, `.lock` suffixes, and names over 1024 bytes.
A CLI Git view name is one non-empty root-name segment. View `origin` is stored under the ordinary root `git/origin`.
## SSH source endpoints
SSH transfer sources use:
```text
ssh://[user@]host[:port]/absolute/repository/path
```
The repository component must decode to one absolute UTF-8 path of at most 16 KiB and may use URL percent escapes. Query strings and fragments are rejected. Bracketed IPv6 literals are accepted. User and host parsing is deliberately strict so no endpoint component becomes shell syntax.
OpenSSH, rather than the URL, supplies identity files, host-key policy, proxy jumps, agents, and other connection settings.
# Local Repository
> Standard persistent profile layout, process coordination, collection behavior, and operational boundaries.
`Repository::local(path).await` opens the standard persistent profile. CLI commands that open a local repository use the same profile. `sync` selects its source and destination explicitly; archive inspection needs no repository.
Opening a new profile creates its payload directory, lock files, and logical database. Opening an already-current profile performs no logical state write. Schema initialization or upgrade may write and is coordinated like any other state mutation.
## Layout
The main paths are:
```text
/
blobs/
blobs/b3//
chunks/b3//
bao/b3//
casita.sqlite
casita.sqlite.online-pins
gc.lock
spill/
runs/
```
| Path | Purpose |
| --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `casita.sqlite` | Revisioned object records, ordered links, named roots, and current repository revision, plus two local accelerators: verified closures and the ingest cache |
| `blobs/blobs/b3/...` | Multi-chunk and empty-payload manifests |
| `blobs/chunks/b3/...` | Zstd-compressed FastCDC chunks |
| `blobs/bao/b3/...` | Optional derived Bao outboards for verified range reads |
| `casita.sqlite.online-pins` | Durable scoped pins and deletion claims |
| `gc.lock` | Cross-process single-collector mutex |
| `spill/` | Temporary traversal state, present only while a traversal exceeds its memory budget |
| `runs/` | Temporary checkouts for `casita run`; an interrupted run may leave files here |
Spill files are not repository state. A traversal that outgrows `SpillLimits` moves its visited set and work queue into a temporary database here, deletes it when the traversal ends however it ends, and holds a lock file meanwhile so a repository open sweeps abandoned files without disturbing a live traversal in another process.
## Spill telemetry
`fsck()` reports traversal temporary-state use in `FsckReport::spill`, and `collect()` reports the same information in `CollectionOutcome::spill`. The `SpillMetrics` summary has two fields:
* `files_opened`: temporary spill databases opened during the operation.
* `peak_bytes`: the greatest aggregate footprint observed for active spill databases, including their SQLite WAL files.
Metrics are captured before cleanup. It is therefore normal for a completed operation to report nonzero values while `spill/` is empty afterwards. They are performance diagnostics, not persistent repository state or a quota guarantee.
The sharding key is lowercase BLAKE3 hex; `\` is its leading shard. Single-chunk payloads omit a separate manifest because their chunk and payload digests are identical.
> **Caution**
>
> This physical layout is an implementation detail of the standard local profile, not a logical format contract. Do not construct object keys from paths, edit the database, replace manifests, or depend on compression and chunking choices in an application protocol.
## Logical state
The SQLite-format database stores four generic concerns:
* one opaque current repository revision;
* immutable object keys, physical payload IDs, and verified sizes;
* canonical ordered forward links; and
* validated named roots selecting exact object keys.
Logical state is authoritative. Casita does not reconstruct records by scanning the payload directory. Physical data without a record is collectible residue; a record with no physical payload is an integrity failure.
## Local accelerators
The same database holds two tables that are not logical state. Neither is consulted by a reader, neither participates in the repository revision, and either may be discarded: losing one costs work, never correctness.
* **Verified closures** record which object graphs have already been checked whole. Records are immutable and only collection removes them, so a closure that verified once stays verified; publication, checkout, and synchronization stop their walk at a remembered one instead of re-reading everything beneath it. File contents need no entry: a raw blob’s record has no links and names its own payload, so its presence is the proof. `fsck` ignores the table, because reading the bytes back is exactly what an audit is for.
* **The ingest cache** records which content each imported file held, keyed by the device, inode, size, and both timestamps the walk observed. See [Imports](../../concepts/imports/).
Both are pruned by collection alongside the objects they refer to, and every digest either one produces is confirmed against committed state before it is used, so an entry that outlives its object costs a lookup rather than returning a stale answer.
The local state engine uses a multi-process WAL. Casita serializes logical writes while independent readers may proceed concurrently. The current pre-release database layout is schema version 6, including object creation generations for online snapshot retention. Existing databases with any other `PRAGMA user_version` are rejected without mutation; development-only schemas from before the first release must be recreated or re-imported. A future schema change will require an explicit offline migration rather than modifying a repository during ordinary open.
Every commit syncs the WAL before it is acknowledged (`synchronous = FULL`). On Linux and Windows that sync reaches stable storage. On macOS it reaches the drive’s volatile cache, so a power loss may discard the most recent acknowledged commits. Deletions that a commit allows (collection sweeps and catalog reclamation) first flush the drive cache (`F_FULLFSYNC`), so a power loss cannot keep those deletions while losing the commit behind them. The repository reopens consistent, at an earlier revision. Casita flushes before every deletion batch, rather than on every commit. WAL size and modification time cannot identify commits that finished syncing after an earlier flush.
## Physical payload behavior
Payload identity is always BLAKE3 over complete plaintext bytes. The default chunk store uses content-defined FastCDC boundaries targeting 256 KiB, bounded at half and twice that average, then compresses chunks with Zstd.
These choices do not affect logical identity:
* equal plaintext payloads share the same `BlobId`;
* equal regions may share physical chunks across different payloads;
* changing chunking, compression, or backend does not change object keys; and
* every chunk is verified after decompression, while a complete sequential read verifies the whole payload at EOF.
Bao outboards are derived physical state. Their presence enables independent verified range reads but is not required to retain or transfer a logical object.
The default store admits at most 64 MiB of plaintext chunk work concurrently across its clones. Transfer independently defaults to a 64 MiB in-flight byte budget. Both are deployment controls rather than durable format parameters.
## Process coordination
Mutation sessions, retention holds, and transfers register durable scoped pins. Readers and writers can continue during collection; their pins retain the logical objects and physical data they use. Collectors take `gc.lock` so only one collection or initialization operation runs across processes.
`collect()` waits for a competing collector, while `try_collect()` returns `Busy`. Operating-system lock release handles collector crashes. Data pins, prune fences, and deletion claims remain durable until their exact ownership is resolved; elapsed time never makes an operation safe to forget.
Deleting or replacing `gc.lock` while Casita processes are running breaks collector coordination.
## Collection order
Collection marks named-root closures and active pin scopes in one immutable snapshot. It then:
1. atomically commits a logical state containing only the marked records;
2. deletes unreferenced payload manifests; and
3. deletes unreferenced chunks.
Logical prune precedes physical deletion. A physical deletion failure may leak space for a later run but does not invalidate reachable logical state.
The standard local profile knows that state and payloads share one filesystem. If the logical prune fails specifically because storage is full, it may delete only the already-marked stale physical set to create emergency headroom, reopen the writer, and retry the prune. The pin ledger separately reserves bounded bookkeeping space for ownership transitions while capacity is available.
Before a local `MutationSession` registers its staging pin, Casita samples disk usage. At the default 80% threshold it attempts nonblocking collection. If usage remains at or above 75%, it releases least recently used evictable roots and vacuums after each release. Permanent roots remain. Completed pressure passes have a 60-second cooldown shared through a repository stamp file. `Busy` admits the mutation and leaves the attempt pending for a later session. Callers may still run `gc`, invoke `collect()` directly, or schedule `DiskPressurePolicy::probe_and_collect`. On a filesystem with nonzero reported capacity, zero free bytes always triggers an attempt when the cooldown permits.
## Deterministic publication crash tests
The library test suite kills a child process at recorded publication steps, then reopens the repository and checks committed roots, payload bytes, and subsequent collection. Run the matrix with:
```sh
devenv shell cargo test --all-features --lib blob::crash_tests -- --nocapture
```
The matrix covers immutable object writes, catalog publication and rebase, SQL root changes, and process death before or after commit. It verifies that visible roots have complete graphs, acknowledged data survives, and a reopened repository can accept new work. It runs in the normal Linux, macOS, and Windows library suites without an external server.
These tests model process death while the operating system remains alive. They do not model power loss, torn sectors, or remote object-store durability. Unix also tests directory flush boundaries.
## Backup and restore
There is currently no separate repository snapshot command. For a simple filesystem backup:
1. stop or quiesce every process that can write, publish, retain, or collect;
2. copy the complete repository root, including the database and payload tree;
3. preserve filesystem metadata needed by the database and regular files; and
4. open the restored copy with Casita and run `fsck` before relying on it.
Copying only `casita.sqlite` loses payloads. Copying only `blobs/` loses the logical keys, links, roots, and revision that make those bytes meaningful. Ordinary copy tools do not participate in Casita’s advisory locks, so do not assume a live file-by-file copy is an atomic snapshot.
Treat an independently restored copy as a separate repository. Revisions are opaque local state tokens and must not be used to order or equate later states across the original and restored repositories.
## Operational commands
| Goal | Command |
| ---------------------------------------------------------------------- | -------------------------------------------------- |
| Inspect roots | `casita --repository PATH root ls` |
| Preview reclaimable data | `casita --repository PATH gc --dry-run` |
| Collect unreachable data | `casita --repository PATH gc` |
| Verify logical and physical integrity, and safely repair when possible | `casita --repository PATH fsck [--source REPLICA]` |
Do not use `gc` as a repair tool for reachable corruption. Preserve the affected repository, inspect the `fsck` findings, and restore or re-import the authoritative data as appropriate.
# Object Formats
> Built-in namespaces, native IDs, and the links each format retains.
An `ObjectKey` names a logical object using a format namespace and native ID. Its `ObjectRecord` also stores a `BlobId`, the BLAKE3 digest of the complete plaintext payload, plus direct links. The two IDs can differ: a Git object’s key uses its Git OID, while its payload has a BLAKE3 `BlobId`.
A format verifier checks the key against the payload, enforces the format’s encoding rules, and reproduces the record’s exact sorted, unique links. Those links decide what a root retains, what sync copies, and what collection keeps.
## Built-in registry
`FormatRegistry::builtin()` contains the formats below. The supported `Repository::memory` and `Repository::local` constructors, and the generic `Repository::new`, use it automatically. To choose an exact set of verifiers, use `FormatRegistry::new(...)` with `casita::experimental::Repository::with_formats`. The registry is immutable, and each namespace can have only one verifier. See [Define a Custom Object Format](../../guides/custom-formats/) for a working example. Registry composition is part of the `experimental` API.
## Filesystem formats
| Namespace | Native ID | Direct links |
| --------------------- | ------------------------------------------------ | ----------------------------------------------------- |
| `casita.blob.v1` | BLAKE3 digest of the plaintext bytes | None |
| `casita.directory.v1` | BLAKE3 digest of the canonical directory payload | Child file blobs and directories; symlinks are inline |
A canonical directory preserves names, file sizes, executable bits, recursive directory counts, and symlink targets. It omits timestamps, ownership, ACLs, and extended attributes.
## Application metadata
| Namespace | Native ID | Direct links |
| ------------------ | --------------------------------------------- | ------------------------------ |
| `casita.linked.v1` | BLAKE3 digest of the canonical linked payload | Its sorted, unique object keys |
`casita.linked.v1` stores a set of links and an opaque application body. The built-in verifier checks the framing and links, not the body’s application meaning. `LinkedObjectFormat::new(namespace)` can use that same framing under an application-owned namespace when explicitly registered.
## IPLD formats
| Namespace | Native ID | Direct links |
| ---------------- | --------------------------------------------------------------------- | -------------------------- |
| `ipld.raw.v1` | Canonical CIDv1 with raw codec `0x55` and BLAKE3-256 multihash | None |
| `ipld.linked.v1` | Canonical CIDv1 with Casita codec `0x300001` and BLAKE3-256 multihash | Its sorted, unique CID set |
`ipld.linked.v1` also uses a fixed Casita payload format. It does not accept arbitrary IPLD codecs. Both IPLD namespaces require CIDv1 and a 32-byte BLAKE3 multihash.
## Native Git formats
Native Git objects store their exact Git body. Their key uses the SHA-1 or SHA-256 OID of Git’s `\ \\0` framing. SHA-1 verification uses collision detection.
| Namespace | Native ID | Direct links |
| -------------------------------------------- | ------------------------------------------- | ------------------------------------------------------------------ |
| `git.sha1.blob.v1`, `git.sha256.blob.v1` | Git blob OID | None |
| `git.sha1.tree.v1`, `git.sha256.tree.v1` | Git tree OID | Blobs, symlink blobs, and subtrees; gitlinks are not generic links |
| `git.sha1.commit.v1`, `git.sha256.commit.v1` | Git commit OID | Tree and parent commits |
| `git.sha1.tag.v1`, `git.sha256.tag.v1` | Annotated tag OID | Exact typed target |
| `git.view.v1` | BLAKE3 digest of the canonical view payload | Direct ref targets and optional exact native pack cache |
A Git view fixes one hash format and an immutable set of direct or symbolic refs. It may name a default ref. Symbolic refs resolve within the view and do not add separate links. The payload also records the exact reachable-object inventory; ordinary object links still retain those objects. A CLI view named `origin` is selected by the root `git/origin`.
## Unknown or unavailable formats
Stored records still carry links, so listing and retention can use them without decoding the payload again. A source can traverse those links for transfer, but a destination needs the format verifier to accept the object. New publication and exact closure validation also require it. Existing roots continue to retain the graph. `fsck` reports an unavailable verifier as `Unchecked`, not as corruption by itself.
# Rust API Reference
> The supported Rust application API for built-in repository workflows.
Import request types and the `Importer` trait from `casita::import`; repository handles and shared data types are exported directly from `casita`. `Repository` is a non-generic handle: backend traits and protocol details do not appear in its signatures. The `native` feature is enabled by default.
The supported surface contains repository handles, importer requests and reports, and portable identity and filesystem types. Implementation modules remain private.
## Repository API
With `native`, the crate exports `Repository`, `Reader`, `VerifiedReader`, `MetadataReader`, `RetainedReader`, `Error`, `ErrorKind`, `CollectionReport`, `IntegrityReport`, `IntegrityIssue`, `IntegrityIssueKind`, and `IntegrityDisposition`.
| Area | `Repository` methods |
| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- |
| Open | `local`, `memory`; `s3` with the experimental S3 storage profile |
| Import | `import` with `BlobImport`, `CopyImport`, `FilesystemImport`, `TarImport`, `CasitarImport`, or `GitImport` / `GitClosureImport` (`git`) |
| Filesystem | `checkout` |
| Objects | `object`, `open`, `open_verified` |
| Consistent reads | `metadata_reader`, `retained_reader` |
| Application metadata | `get`, `scan`, `commit` |
| Names | `root`, `roots`, `set_root`, `compare_and_set_root`, `remove_root`, `root_retention`, `set_root_retention`, `set_root_with_retention`, `touch_root` |
| Archives | `export_casitar` |
| Maintenance | `preview_collection`, `collect`, `try_collect`, `vacuum`, `fsck`, `flush` |
`Importer\` is the sole public import contract for both repository handles. It consumes an input request and returns that importer’s associated `Report` and `Error`. The repository’s `import(input)` method delegates to this trait; there are no public format-specific import methods. Built-in requests use application `Error` with the standard handle and typed engine errors with experimental compositions. Custom importers can implement the same trait.
`GitClosureImport::new(objects_dir, roots)` imports type-qualified native Git roots directly from an object directory. It follows packs and alternates and reuses complete stored subtrees across unrelated revisions and repositories. It creates no named roots or serving-view inventory. The returned `GitClosureImportOutcome` contains a `report` and a retained `reader`; keep the reader alive until application roots have been published. A fully stored selection does not access the source directory. A present blob is complete on its own, so imports store a completeness record for a blob only when it is a selected root. Completeness records are committed in bounded batches, and only once the whole selection is proven complete; repeating an interrupted import completes them. Report counters describe work performed and reuse boundaries, not the size of the complete reachable graph.
`Reader` implements Tokio `AsyncRead` and `AsyncSeek`. It keeps the selected object’s content protected from collection until dropped. `VerifiedReader` provides sequential reads that authenticate bytes before returning them. Both expose their `ObjectRecord` through `record()`.
`metadata_reader` keeps one metadata revision stable but does not protect payloads from collection. Use `retained_reader` when a root or metadata lookup and the content read must share a protected snapshot. `get`, `scan`, and `commit` operate on namespaced application metadata where the backend supports it. `Error` exposes `kind()` and `retry_disposition()` and preserves the standard error source chain.
`set_root` unconditionally creates or replaces a name after verifying the complete target graph. `compare_and_set_root(name, expected, target)` publishes only if the name is absent (`None`) or still points to the expected key (`Some(&key)`). It returns `false` on mismatch and retries unrelated revision conflicts internally. It compares current values, not root-change history, so a delayed duplicate create can succeed after a removal; the [multi-owner guide](/guides/s3-multi-owner/) shows a fenced transition built from `commit`. `remove_root` likewise requires an exact expected target. Local roots are permanent by default. `set_root_retention` marks an existing root `RootRetention::Evictable` or `RootRetention::Permanent`; `set_root_with_retention` sets a root and its policy atomically. Applications using an evictable root can call `touch_root` after a successful read to update its eviction order. The CLI’s artifact checkout, restore, and run paths do this automatically. The S3 profile does not support evictable root policy. Filesystem and tar import requests also accept `with_retention`, which publishes the imported root and its policy atomically.
`import(CopyImport::new(source, source_name, destination_name))` copies a complete graph from one retained source snapshot and atomically creates or replaces the destination root, including an occupied destination name.
Cloning a repository shares coordination and caches. Drop active readers before collection or the final `flush()`, and await that flush before runtime shutdown. It finishes payload writes, dropped-lease cleanup, and transient metadata compaction. Local SQLite checkpoints report `ErrorKind::Busy` when a live snapshot prevents truncation; release the snapshot and retry. Checkpointing does not delete payload packs or catalog roots.
`CollectionReport` reports logical-object, physical-payload, and chunk counts. `IntegrityReport` exposes inspected counts, a revision, and `IntegrityIssue` findings categorized by `IntegrityDisposition` and `IntegrityIssueKind`. `is_healthy()` checks for reachable corruption; `is_clean()` requires no findings of any kind.
## Portable data model
These types remain public with `default-features = false`:
* `Digest`, `BlobId`, `DirectoryId`, `ObjectId`, and `DigestError`;
* `Directory`, `Node`, `DirectoryError`, and `DirectoryDecodeError`;
* `PathComponent`, `PathComponentError`, `SymlinkTarget`, and `SymlinkTargetError`;
* `NamespaceId`, `ObjectKey`, `ObjectRecord`, `RootName`, `RootRecord`, `RepositoryRevision`, and `RepositoryGeneration`;
* `NamespaceIdError`, `ObjectKeyError`, `RootNameError`, and `RepositoryRevisionError`; and
* `RetryDisposition`.
Physical chunk identities, backend implementations, format verifiers, and wire framing are not application exports.
`ObjectRecord` and `RootRecord` expose read-only getters. Record construction and binary encoding of keys and records use experimental free functions. `ObjectKey` construction and text parsing/display remain supported, as do `Directory` construction and encoding. The binary formats remain unchanged.
## Experimental boundary
Enable `experimental` and import from `casita::experimental` for the generic `Repository\`, backend and format traits, sessions, detailed options, archive framing, transfer protocols, Git services, Gix adapters, verified range primitives, and backend conformance helpers. Dependency re-exports also live there. These APIs may change between releases.
An existing custom composition migrates to:
```rust
use casita::experimental::{MemoryBlobStore, MemoryMetadataStore, Repository};
let repository = Repository::new(MemoryBlobStore::new(), MemoryMetadataStore::new()?);
```
Both handles use the same storage engine and publication rules. The CLI enables `experimental` for advanced commands. `fuzzing` adds parser adapters solely for test harnesses.
See the [Library guide](../../library/) for supported workflows or the [Experimental Rust API](../experimental-rust-api/) for advanced contracts. `cargo doc --no-deps` shows the default public API; `--all-features` includes the experimental namespace and optional integrations.