Compilation is expensive. A medium-sized C++ project might take minutes to build from scratch. When you're iterating on code, most of that time is wasted recompiling modules that haven't changed. cmod's artifact cache eliminates this waste through content-addressed storage.
Content-Addressed Storage Explained
In a content-addressed system, each artifact is stored under a key derived from its contents and inputs. If the inputs haven't changed, the key is the same, and the cached artifact is returned instead of recompiling.
This is fundamentally different from timestamp-based caching (like Make uses). Timestamps can lie — touching a file changes its timestamp without changing its content, and clock skew between machines makes timestamps unreliable. Content hashes are deterministic and portable.
What Goes Into a Cache Key
cmod computes cache keys using SHA-256 hashes of:
- Source file content: The SHA-256 hash of the source file itself
- Compiler identity: Compiler path, version string, and target triple
- Compilation flags: Optimization level, warnings, feature flags
- Dependency BMI hashes: The cache keys of all imported modules
- C++ standard: c++20, c++23, etc.
// Conceptual cache key computation
cache_key = SHA256(
source_content_hash,
compiler_version,
target_triple,
optimization_level,
cpp_standard,
sorted(dependency_bmi_hashes),
)
The dependency BMI hashes are critical. If module A imports module B, and B's source changes, then B's cache key changes, which changes A's cache key (because it includes B's BMI hash). This cascading invalidation is automatic and precise — only modules actually affected by a change are recompiled.
Cache Structure
The local cache lives at ~/.cache/cmod/ (or $XDG_CACHE_HOME/cmod/) and is organized by hash prefix for efficient lookup:
~/.cache/cmod/
├── artifacts/
│ ├── 3f/
│ │ └── 3f7c4d8a1b... # BMI or object file
│ ├── a0/
│ │ └── a0b8a4e19f...
│ └── ...
└── metadata/
└── cache.db # Index for fast lookups
Cache Operations
# See cache status
cmod cache status
Cache location: ~/.cache/cmod
Total size: 142 MB
Entries: 847
Hit rate: 94.2% (last session)
# Remove old entries (keep last 30 days)
cmod cache gc
# Clear everything
cmod cache clean
# Build without cache (for verification)
cmod build --no-cache
Why This Matters for Incremental Builds
Consider a project with 50 modules. You change one leaf module. With content-addressed caching:
- The changed module is recompiled (cache miss)
- Its direct dependents are recompiled (their cache keys changed because the dependency hash changed)
- Everything else is a cache hit
In practice, this means a typical edit-compile-test cycle touches 2-5 modules instead of 50. The build goes from 30 seconds to 2 seconds.
Cache Correctness Guarantees
A cache is only useful if it's correct. cmod's content-addressed approach provides strong guarantees:
- No false hits: If the cache key matches, the output is guaranteed correct. The key includes every input that affects the output.
- No stale entries: There's no concept of "stale" in a content-addressed cache. An entry is either the right one (key matches) or doesn't exist.
- Safe concurrent access: Cache writes use atomic file operations (write to temp file, then rename). Multiple cmod instances can share the same cache safely.
- Portable across machines: Because keys are content-based (not path-based), cache entries are valid on any machine with the same compiler and configuration.
The Path to Distributed Caching
Content-addressed caching naturally extends to distributed (remote) caching. Because cache keys are deterministic and machine-independent, a cache entry produced on one machine is valid on any machine with the same compiler setup.
Phase 3 of cmod's roadmap includes remote cache support:
# Push local cache to remote (planned)
cmod cache push --remote s3://my-cache-bucket
# Pull from remote cache (planned)
cmod cache pull --remote s3://my-cache-bucket
# Build with remote cache enabled (planned)
cmod build --cache-remote s3://my-cache-bucket
The workflow: when a developer builds locally, cache entries are pushed to a shared remote cache. When a CI job or another developer builds the same code, they get cache hits from the remote cache instead of recompiling. This is the same principle behind Bazel's remote execution and sccache, but integrated directly into the build tool.
Comparison with Other Caching Approaches
vs. ccache/sccache
ccache and sccache cache individual compiler invocations. They work at the file level and don't understand module dependencies. cmod's cache is module-aware — it knows that changing module B invalidates module A's cache, even if A's source hasn't changed. This is more precise and correct.
vs. Bazel's remote cache
Bazel's approach is similar in principle (content-addressed, deterministic keys), but requires Bazel's entire build system. cmod provides the same caching benefits with a much simpler tool and configuration.
vs. precompiled headers (PCH)
PCH are a compiler-specific optimization that predates modules. They're fragile (sensitive to include order and macro state), non-portable, and don't compose well. Module BMIs are a language-level feature with well-defined semantics.
Practical Tips
- Commit your lockfile: This ensures all team members resolve the same dependency versions, maximizing cache hit rates
- Standardize compiler versions: Different compiler versions produce different cache keys. Pin your toolchain in
cmod.toml - Use the same optimization flags: Debug and release builds have different cache keys (by design). Don't mix them
- Run
cmod cache gcperiodically: Remove entries older than 30 days to keep cache size manageable
cmod's content-addressed cache turns C++ compilation from a slow, repetitive process into a fast, incremental one. Combined with C++20 modules and deterministic builds, it makes C++ development feel as responsive as working in a language with a REPL.
Get started with cmod and see the cache in action.