Lumen
Operate

Project archives

Export and recover portable semantic project records without leaking operational credentials.

Outcome

A project archive preserves notebooks, normalized executions and outputs, artifacts, claims, evidence, validations, and captured variables in a versioned format. Import remaps every portable identity into the target project, records the source-to-target map, and keeps the source provenance and omission disclosures queryable in PostgreSQL.

The archive is a recovery and transfer format. It is not a database backup and it does not claim that omitted proprietary inputs can be reproduced.

Prerequisites

Use an authenticated actor who can read the source project. Before import, authorize that actor for the exact target workspace and project, choose an empty target when portable semantic names could conflict, and make the canonical artifact store available to the trusted worker. No public import API or browser-side database restore exists.

Steps

  1. Export the latest project archive through the authenticated project endpoint.
  2. Preserve the original bytes and verify the manifest, database digest, artifact inventory, and omissions before transfer.
  3. Submit the bounded stream, target scope, and actor to the worker-owned import operation.
  4. Inspect the saved archive digest, identity map, provenance transformations, and omissions after the transaction commits.

Export

Call GET /api/projects/{project_id}/archives/latest with the same bearer authorization used for the project. The response media type is application/vnd.lumen.project-archive+zip and the browser must treat it as a download, not inspect it as executable content.

The ZIP uses stored regular-file members only:

  • manifest.json declares format version 1, source scope, export time, database digest, artifact inventory, provenance, and omissions;
  • database.json contains the exact version-1 semantic table inventory; and
  • each artifacts/art_…/<sha256> member contains verified immutable bytes for one artifact record.

The generated backend limits publish the current archive, manifest, database, and member ceilings directly from the owning source assignments.

Export runs in a repeatable-read project snapshot. Every artifact is verified against its canonical digest, size, media type, and tenant storage key while it is streamed. A missing or corrupt object fails the export instead of producing a plausible incomplete archive.

Import trust boundary

Import is a worker-owned operation. The caller first authorizes the human actor for the exact target workspace and project; the worker then receives the bounded archive stream, target scope, actor, and the configured canonical artifact store. The public app process does not receive worker database or object-store write credentials, and imported bytes never pass through a sandbox.

Before any semantic row is committed, import verifies the ZIP member inventory, rejects duplicate JSON keys and compressed or encrypted members, checks the database digest and byte limit, and streams every artifact through the normal trusted finalizer. PostgreSQL insertion happens in one transaction after those checks. A digest mismatch, broken reference, target-name conflict, or failed constraint rolls back the import; immutable bytes finalized before a database failure may remain unreferenced and are safe to reclaim by normal retention tooling.

Repeating the same archive digest against the same target returns the existing import record. It does not duplicate documents, claims, evidence, or events.

Provenance transformations

Portable semantic objects retain their wording, revisions, evidence roles, validation outcomes, artifact digests, and historical validity/currentness. New opaque IDs replace source IDs everywhere, including structured metadata and evidence locators. The import record stores the complete identity map.

Operational lineage cannot be copied as if it ran inside the target deployment:

  • executions become imported executions with no live run, attempt, runtime session, or requested environment;
  • claim and block authors become the import actor while their original lineage remains in archive provenance;
  • variable values become explicitly imported values and preserve their original execution, block, environment, input, and content hash in archive_source_provenance; and
  • provider jobs, workflow history, sandboxes, approvals, secrets, connector grants, capability tokens, and provider credentials remain omitted.

Verify

After import, verify the saved archive digest and source scope, then open an imported claim's bounded evidence neighborhood. The claim revision, evidence reference, variable value, execution output, and artifact must all use remapped target identities. Opening the artifact must return bytes whose digest matches the imported manifest.

Inspect every omission. reproducibility_impact: blocks means an operator must restore or reauthorize the missing input before describing the imported result as reproducible.

Recover

If import fails, keep the rejected archive and error code for diagnosis; do not bypass digest or reference checks. Correct the source archive or choose an empty target when semantic names conflict, then retry the original bytes. The same valid digest is idempotent.

Project archives complement PostgreSQL backups and object-store recovery. Use the backup and restore runbook for disaster recovery of the deployment itself.

Next task

Prepare deployment backup and restore.

On this page