Omnigraph

Cluster configuration reference

Cluster commands read a directory containing cluster.yaml:

Cluster commands read a directory containing cluster.yaml:

omnigraph cluster validate --config ./company-brain

--config defaults to the current directory. Unknown fields and duplicate YAML keys are errors so that misspelled intent is never ignored.

Complete shape

version: 1

metadata:
  name: company-brain

# Omit for a cluster stored in this directory.
storage: s3://company-data/omnigraph/company-brain

state:
  backend: cluster
  lock: true

providers:
  embedding:
    default:
      kind: openai-compatible
      base_url: https://api.example.com/v1
      model: text-embedding-3-large
      api_key: ${EMBEDDING_API_KEY}

graphs:
  knowledge:
    schema: knowledge.pg
    queries: queries/
    embedding_provider: default
    external_blobs:
      allow:
        - base: s3://company-assets/knowledge/
          scope: server_safe

policies:
  graph-access:
    file: graph.policy.yaml
    applies_to: [knowledge]
  server-access:
    file: server.policy.yaml
    applies_to: [cluster]

Top-level fields

FieldRequiredMeaning
versionyesConfiguration schema; currently 1
metadata.namenoDisplay name
storagenoCluster root; direct default is the config directory, remote apply defaults to the selected server
state.backendnoOmit or set to cluster
state.locknoExclusive writer admission; omit or set to true; false refuses execution
providers.embeddingnoNamed embedding provider profiles
graphsnoGraph declarations keyed by graph ID
policiesnoPolicy bundles keyed by bundle name

Credentials are process configuration and must not appear in cluster.yaml.

Graphs

Each graph requires a schema file:

graphs:
  knowledge:
    schema: knowledge.pg

Optional fields:

FieldMeaning
queriesStored-query files, directories, or explicit name mappings
embedding_providerName under providers.embedding
external_blobsAllow-list for new external Blob references

Query declarations support three forms:

# Every declaration in top-level *.gq files in a directory
queries: queries/

# Every declaration in these files or directories
queries: [people.gq, reports/]

# Explicit registry names
queries:
  find_experts:
    file: knowledge.gq

Unreadable files, parse errors, duplicate query names, and queries that do not type-check against the graph's desired schema fail validation.

Embedding providers

Provider kind may be openai-compatible, openai, gemini, or mock. Real providers require api_key: ${ENVIRONMENT_VARIABLE}; inline secrets are rejected. The serving process resolves the environment variable at boot and during live deployment preflight for affected graphs. cluster validate, plan, and direct apply do not resolve it. Vector dimensions remain part of the graph schema.

See Embeddings for provider behavior.

External Blob references

New external references are denied unless their normalized URI falls under an allowed base:

external_blobs:
  allow:
    - base: s3://company-assets/knowledge/
      scope: server_safe

server_safe permits the base for a served graph. embedded_only is for an embedded host and may permit a local file:// directory; it is not installed by the HTTP server or direct-store CLI. Bases must be absolute, non-overlapping, and free of credentials, query strings, fragments, and path traversal.

A base must also name storage outside the cluster's storage root: the config directory when storage is omitted, or the storage URI. That root holds every graph and the applied state, and ingress reads with the process's own storage credentials, so a base over it would let any writer copy another graph's data or the cluster ledger into a readable Blob value. cluster validate, plan, and apply refuse such a base with external_blob_base_overlaps_storage_root, whatever its scope. Put external objects under a sibling prefix instead, for example s3://company-assets/cluster-external/ beside storage: s3://company-assets/cluster. A server that finds an overlapping server_safe base in the applied state quarantines that graph and serves the others; if no applied graph is left to serve, startup fails with cluster_no_healthy_graphs. An embedded handle refuses a policy whose base overlaps its own graph root.

A base is compared only with a storage root of its own kind: an s3:// base with an s3:// root, a file:// base with a local root. When the root is spelled with a path component a base URI cannot express (an empty component such as s3://bucket/a//cluster, or a percent sign in a local path), a same-kind base is refused with external_blob_storage_root_uncomparable, because disjointness cannot be proven. Moving the base does not clear that code; the storage root spelling does.

The allow-list controls which external objects an authorized writer may cause the process to inspect. Cedar policy separately decides who may write. See Blob values.

Policies

policies:
  graph-access:
    file: graph.policy.yaml
    applies_to: [knowledge, catalog]
  registry-access:
    file: server.policy.yaml
    applies_to: [cluster]

A bundle targets either graph IDs or the cluster server scope, never both. Only one bundle may bind a given graph or the cluster scope. See Authorization.

Storage

For direct apply, omitted storage puts applied state and graph data under the config directory. cluster apply --server instead uses the selected server’s canonical root when storage is omitted; an explicit absolute path or file://, s3:// or az:// URI must match it. Remote apply rejects relative storage paths and reads only the source bundle from the caller’s filesystem.

Use the standard storage credential environment for the chosen backend. Azure is a qualification preview and requires the admission wrapper for every writer; see Deployment.

Declared paths

Every schema, queries, and policy file path is resolved against the directory that holds cluster.yaml. A relative path must stay inside that directory: a .. segment is refused with config_path_escape, and a path that reaches its file through a symbolic link, on the way to it or as a query file discovered inside a declared directory, is refused with config_path_symlink. Each diagnostic names the setting that declared the path. The bundle is one directory of files read exactly as declared, so what gets applied never depends on something outside it.

An absolute path is accepted as given and is not checked for either shape. Prefer relative paths; they are what keep a bundle portable and hermetic.

Command behavior

CommandChanges graph or cluster state?Use
validatenoParse and type-check the declaration
plannoPreview desired differences; unsupported deployment effects still refuse
applyyesConverge graph inventory, schemas, queries, policies and provider/Blob settings
statusnoRead recorded deployment and lock status; server status also observes activation
observenoReport current observations without changing authority
force-unlockyesRemove one proven-stale lock by exact ID
upgrade-ledgerstate onlyConvert a stopped legacy ledger without resetting graph data

All execution uses ledger v2 and requires state.lock: true. With --server, apply submits to the existing writer and activates without restart. Direct apply owns exclusive admission and requires an explicit handoff before serving. Retained graph roots and storage formats stay fixed. Policy/provider/Blob bindings can change; removing a graph declaration deletes its managed storage and history. See deployment boundaries. Apply does not load rows or start servers. Root-addressed status, reconciliation and conversion do not use cluster.yaml; see the deployment workflow.

Limits

Configuration loading refuses limits before schema effects. Shared files with identical bytes count once toward the aggregate source limit.

ResourceLimit
Each cluster.yaml, schema, query or policy source1 MiB
Distinct source bytes in one captured bundle8 MiB
Declared resources (each graph and schema count separately)4096
Query discovery paths plus directory entries, including ignored files4096
Encoded immutable deployment bundle16 MiB
Encoded cluster ledger16 MiB
Lock metadata read64 KiB
Outstanding deployments1
Retained deployment results32 records, 4 MiB total, 1 MiB each
Deployment actor / resource address / canonical root256 / 512 / 4096 UTF-8 bytes

Deployment admission checks the actual encoded prepared intents and reserves ledger/result space to record completion before accepting schema effects. JSON escaping and prepared state can make a deployment exceed its encoded limit even when raw source bytes fit. Result history may evict older completed records to admit a new deployment; outstanding authority is retained. These bounds do not cap graph data, native manifest-history scans or total process memory.

See Operating a cluster for the end-to-end workflow.

On this page