Project archives
Export and recover portable semantic project records without leaking operational credentials.
Outcome
A project archive preserves notebooks, normalized executions and outputs, artifacts, claims, evidence, validations, and captured variables in a versioned format. Import remaps every portable identity into the target project, records the source-to-target map, and keeps the source provenance and omission disclosures queryable in PostgreSQL.
The archive is a recovery and transfer format. It is not a database backup and it does not claim that omitted proprietary inputs can be reproduced.
Prerequisites
Use an authenticated actor who can read the source project. Before import, authorize that actor for the exact target workspace and project, choose an empty target when portable semantic names could conflict, and make the canonical artifact store available to the trusted worker. No public import API or browser-side database restore exists.
Steps
- Export the latest project archive through the authenticated project endpoint.
- Preserve the original bytes and verify the manifest, database digest, artifact inventory, and omissions before transfer.
- Submit the bounded stream, target scope, and actor to the worker-owned import operation.
- Inspect the saved archive digest, identity map, provenance transformations, and omissions after the transaction commits.
Export
Call GET /api/projects/{project_id}/archives/latest with the same bearer authorization used for
the project. The response media type is application/vnd.lumen.project-archive+zip and the browser
must treat it as a download, not inspect it as executable content.
The ZIP uses stored regular-file members only:
manifest.jsondeclares format version1, source scope, export time, database digest, artifact inventory, provenance, and omissions;database.jsoncontains the exact version-1 semantic table inventory; and- each
artifacts/art_…/<sha256>member contains verified immutable bytes for one artifact record.
The generated backend limits publish the current archive, manifest, database, and member ceilings directly from the owning source assignments.
Export runs in a repeatable-read project snapshot. Every artifact is verified against its canonical digest, size, media type, and tenant storage key while it is streamed. A missing or corrupt object fails the export instead of producing a plausible incomplete archive.
Import trust boundary
Import is a worker-owned operation. The caller first authorizes the human actor for the exact target
workspace and project; the worker then receives the bounded archive stream, target scope, actor, and
the configured canonical artifact store. The public app process does not receive worker database or
object-store write credentials, and imported bytes never pass through a sandbox.
Before any semantic row is committed, import verifies the ZIP member inventory, rejects duplicate JSON keys and compressed or encrypted members, checks the database digest and byte limit, and streams every artifact through the normal trusted finalizer. PostgreSQL insertion happens in one transaction after those checks. A digest mismatch, broken reference, target-name conflict, or failed constraint rolls back the import; immutable bytes finalized before a database failure may remain unreferenced and are safe to reclaim by normal retention tooling.
Repeating the same archive digest against the same target returns the existing import record. It does not duplicate documents, claims, evidence, or events.
Provenance transformations
Portable semantic objects retain their wording, revisions, evidence roles, validation outcomes, artifact digests, and historical validity/currentness. New opaque IDs replace source IDs everywhere, including structured metadata and evidence locators. The import record stores the complete identity map.
Operational lineage cannot be copied as if it ran inside the target deployment:
- executions become imported executions with no live run, attempt, runtime session, or requested environment;
- claim and block authors become the import actor while their original lineage remains in archive provenance;
- variable values become explicitly imported values and preserve their original execution, block,
environment, input, and content hash in
archive_source_provenance; and - provider jobs, workflow history, sandboxes, approvals, secrets, connector grants, capability tokens, and provider credentials remain omitted.
Verify
After import, verify the saved archive digest and source scope, then open an imported claim's bounded evidence neighborhood. The claim revision, evidence reference, variable value, execution output, and artifact must all use remapped target identities. Opening the artifact must return bytes whose digest matches the imported manifest.
Inspect every omission. reproducibility_impact: blocks means an operator must restore or reauthorize
the missing input before describing the imported result as reproducible.
Recover
If import fails, keep the rejected archive and error code for diagnosis; do not bypass digest or reference checks. Correct the source archive or choose an empty target when semantic names conflict, then retry the original bytes. The same valid digest is idempotent.
Project archives complement PostgreSQL backups and object-store recovery. Use the backup and restore runbook for disaster recovery of the deployment itself.