A service which exports main metadata of many blockchain networks as Prometheus metrics, and publishes per-endpoint Chain-State Oracle observations (ADR-190).
poetry installcp ./config/config.example.yaml ./config/config.yamlpython3 -m aiohttp.web -H 0.0.0.0 -P 8000 src.main:init Tests: poetry run pytest.
Network-level (unchanged, read by chainy's head-lag check — monotonic high-water mark across all sources):
blockchain_network_block_number{network} - The current block height of the blockchainblockchain_network_block_timestamp{network} - The timestamp of the last block height updatePer-endpoint oracle attribution (hashes never appear in metrics; they travel over NATS only):
oracle_endpoint_client_info{chain,endpoint_id,source,provider_id,client_type} - always 1; client implementation and attribution of each observed endpointoracle_endpoint_client_build_info{chain,endpoint_id,source,provider_id,client_type,client_build,client_version} - always 1; the build and version each observed endpoint runs, derived from its web3_clientVersion by the contract's parse_build_info (conduit-op-reth/v2.2.3-38f6b93/… gives conduit-op-reth, 2.2.3-38f6b93), at most 64 bytes each, unknown when nothing qualifies. client_type stays the diversity signal: a vendor build is an alias of its client, so the build is visible here without counting as a second implementation. The raw string stays on NATS. Replaced on upgrade, withdrawn with client_infooracle_endpoint_head_lag_blocks{chain,endpoint_id,source,provider_id} - blocks the endpoint's latest reading sits behind the highest latest reading among the chain's current sources (sources excluded from observations do not count; with a single source it is 0). Set once per scrape round, after every source has reported, so all sources of a chain are measured against the same reference. Sources are read at slightly different moments, so on fast chains it carries a few blocks of noise; alert on min_over_time(…[1m]), which a source that is really behind never brings to 0oracle_endpoint_head_age_seconds{chain,endpoint_id,source,provider_id} - age of the endpoint's latest block when it was read (read time minus block timestamp). It needs no peers, so it also catches a sole source that answers but is stuck, and it is in seconds on every chain. Lag against peers is a query, free of read-time skew and of pauses in block production (a quiet chain ages every source alike):
oracle_endpoint_head_age_seconds - on (chain) group_left () min by (chain) (oracle_endpoint_head_age_seconds).
Block timestamps have 1s resolution; chains that only produce blocks on demand (Avalanche C-chain, quiet testnets) age while healthy, so alert on the relative form there. The absolute form trusts the exporter's clockoracle_endpoint_finality_delay_seconds{chain,endpoint_id,source,provider_id,tag} - how far the endpoint's safe / finalized head sits behind its latest, from block timestamps of its last safe/finalized fetch (only while publishing). One that keeps growing is a source whose finality has stopped; across a chain's sources it is the measured input for the aggregator's per-network settle depth (CORE-15327). Withdrawn with the other head seriesoracle_observations_total{chain,result} - outcome per scrape: published, stored (no NATS configured), publish_failed, unknown_chain (label not in the chain registry), no_hash (protocol without hash capture yet), no_chain_id, chain_id_mismatch, invalidchain on every oracle and per-source metric is the chain-state-oracle registry slug (see below), not the
network label of the legacy gauges; while a label is unresolved it falls back to the label itself.
Per-source scrape health (CORE-15586) — effective participation, not configured count:
blockchain_network_source_scrapes_total{chain,endpoint_id,source,provider_id,result} - success or a failure class: dns, tls, connect, timeout, network, http_429, http_4xx, http_5xx, rpc_error, parse, unsupported (no scraper for the label, see Configuration), otherblockchain_network_source_last_success_timestamp_seconds{…} - unix time of the last successful scrape, 0 if neverblockchain_network_source_up{…} - 1 while the latest scrape succeeded and is fresh, 0 when failing or staleblockchain_network_source_excluded{…,reason} - 1 while the source scrapes fine but is kept out of observations:
chain_id_mismatch, no_chain_id (the identity probe has not yielded a chainId yet) or invalid. Absent
otherwise, and cleared as soon as a scrape fails or goes stale, so it is only ever set where up is 1.
Sources of chains that emit no observations at all (not in the registry, or no hash capture yet) are never
marked: they contribute to the gaugesup is scrape health; up minus excluded is what actually contributes. The or … * 0 keeps a chain with
no excluded source in the result instead of dropping it from the subtraction:
sum by (chain) (blockchain_network_source_up) # sources scraping successfully
sum by (chain) (blockchain_network_source_up)
- sum by (chain) (blockchain_network_source_excluded or blockchain_network_source_up * 0)
# sources contributing right now (0, not absent, when none)
blockchain_network_source_excluded == 1 # which sources are kept out, and why
time() - blockchain_network_source_last_success_timestamp_seconds # seconds since each source last worked
rate(blockchain_network_source_scrapes_total{result!="success"}[5m]) # failure rate by class
A source with no success in three scrape rounds in a row (or for oracle.stale_after_seconds, when set) is
dropped from /observations and loses its oracle_endpoint_client_info series, so a dead source stops looking
like a quorum member; its counters and last-success timestamp remain. Rounds rather than seconds by default:
a round lasts as long as its slowest endpoint, so a fixed TTL would count one slow round as several missed ones. A source that still scrapes but whose scrape
yields no observation (chain_id_mismatch, no_hash, invalid, …) is withdrawn the same way at once (and, for
chain_id_mismatch, no_chain_id and invalid, marked on blockchain_network_source_excluded), and a
chain_id_mismatch is re-probed on identity_retry_interval so a source fixed upstream returns within a minute. The legacy blockchain_network_* gauges
keep their high-water-mark semantics until chainy moves off them (CORE-15337).
An EVM probe sends the single eth_getBlockByNumber(latest) call this exporter has always made, and the
latest block's hash comes with it. Only where observations are published (oracle.nats_url set) does it
also fetch safe and finalized, batched with latest. An exporter that does not publish sends exactly
what it sent before the oracle.
How often is measured, not configured per chain. Each fetch measures the chain's finality delay: how far
its safe or finalized head sits behind latest (whichever is closer), from the block timestamps of that
one batch, and the smallest across the chain's sources, so one stuck source cannot slow its chain down. The
next fetch comes a quarter of that delay later, which keeps the reported heads within a quarter of where the
chain's own finality already puts them, bounded by oracle.finality_interval_min (60s) and
oracle.finality_interval_max (10 min). Measured on 2026-09-24 against public endpoints of the chains in the
prod config: instant-finality chains (BSC, Polygon, Avalanche, Sonic) sit at the 60s floor, OP-stack L2s at
60s–2 min (their safe moves every few minutes), Ethereum and Arbitrum at about 2 min, and zk rollups
(Linea, zkSync, Scroll) at the 10 min ceiling. A
source that answers neither tag waits the ceiling instead of paying for a useless batch every minute. Each
source's first fetch falls at a random point within the floor, so the batches spread over the rounds.
A batch costs two calls more than the single call. A provider that rejects it gets the single call instead,
and only that call for an hour. Rejecting means a 4xx other than 429, a reply that is not a batch answer, or
three batches in a row that came back with nothing but errors while the single call worked. A batch that
fails (429, 5xx, timeout) never fails the scrape: the batch exists for the oracle only, so the gauges hang on
the single call, which the provider then gets alone for ten minutes, the pre-oracle load on a provider that
is rate-limiting or down, while safe/finalized come from the chain's other sources.
Each probe emits one
observation per endpoint — {number, hash} for latest, and for safe/finalized on the scrapes that
fetched them (between fetches they are absent, never repeated stale), tagged source, providerId,
clientType, chainId — encoded with the canonical JSON profile of
chain-state-oracle (schemaVersion 1.1.0).
oracle.<chain>.obs when oracle.nats_url is set. Publishing is fail-open:
a NATS outage is counted and logged, and never affects the gauges. The connection is made in the background
(oracle.connect_timeout per attempt, retried every 10s), so an unreachable NATS never delays a scrape; the
warm-up scrape waits for the first attempt to settle, one bounded attempt at most, so a healthy deploy does not
count its first round as publish_failed.GET /observations[?chain=<slug or alias>] as a JSON array of the last payload per
endpoint, byte-identical to what goes on the wire (?network= is accepted as a synonym).blockchain_network_source_scrapes_total.web3_clientVersion, eth_chainId) is probed once per identity_probe_interval,
not per scrape, because aggregators rate-limit extra calls; each source's first re-probe falls at a random
point in the second half of the interval, so a fleet probed in one warm-up round does not re-probe in one
round every hour. A method the provider refuses (4xx) is logged with its status and left out of the identity;
without eth_chainId the source is excluded as no_chain_id and re-probed on identity_retry_interval.
Normalization mirrors the contract's
client-identity table (schema 1.1.0 covers the served fleet: nitro, sonic, kaia, avalanchego, …).
Proxy rewriters normalize to unknown. client_type in config overrides the probe (read on every
scrape, so a reload applies it at once, without waiting for the next probe); it is needed
for reth-derived distributions that self-report reth (base-reth, nanoreth, Plasma) and for older
op-reth builds. Aggregators such as dRPC answer with a synthetic string (Geth/v10.0.0/drpc) that
would classify as geth; pin them to client_type: unknown so they never count toward client diversity.eth_chainId differs from the registry's pin for its chain is excluded from observations.Conformance against the shared golden vectors is tested in tests/test_observation.py; see
tests/vectors/SOURCE.md to refresh them.
The chain an observation carries, and the oracle.<chain>.obs subject it goes to, is a slug from the
chain-state-oracle chain registry (src/oracle/chains.json, generated upstream from the platform catalog,
chainstack/core networks.csv): one slug per Chainstack internal network name with the chain id (or Solana
genesis hash) it pins. Producers never invent a subject, so own and external observations for one chain always
meet on one subject even when their Prometheus labels differ (polygon-mainnet is a registry alias of
polygon-pos-mainnet; kaia-mainnet of klaytn-mainnet).
Each endpoint's chain is resolved at config load, in this order: an explicit chain: (must be a registry slug
or alias, otherwise the config is rejected), then additional_network_label, then network. A label the
registry does not know is logged at load, keeps feeding the blockchain_network_* gauges, and is counted as
unknown_chain instead of emitting an observation. The fix is a catalog row in chainstack/core followed by a
registry refresh upstream, or chain: pointing at the slug the label means. The legacy gauges keep the
network label unchanged either way.
src/oracle/registry.py mirrors the Go and Rust lookup semantics and passes the same shared fixture
(tests/vectors/registry/resolve_cases.json, tests/test_registry.py).
oracle:
nats_url: "nats://jetstream-nats:4222" # omit to disable publishing
finality_interval_min: 60 # bounds of the measured safe/finalized fetch interval, seconds;
finality_interval_max: 600 # only while publishing
endpoints:
- network: ethereum-mainnet
url: https://lb.drpc.org/ogrpc?network=ethereum&dkey=...
chain: ethereum-mainnet # registry slug or alias; resolved from additional_network_label / network when omitted
source: external # inferred from host when omitted (*.chainstack.com, *.p2pify.com, *.chainstacklabs.com → own)
provider_id: drpc # inferred from the host's registrable name when omitted (rpc.example.co.uk → example);
# IP and cluster-internal hosts infer `unknown` (set it)
endpoint_id: lb.drpc.org # defaults to the host; unique per chain when set; never put credentials here, it goes on the wire
client_type: op-reth # optional override of the probed client; must be a clientType enum value
The config file is re-read when its mtime changes, so a provider can be added or dropped without an image
deploy. The Helm chart mounts the config Secret as a directory (not a subPath), which Kubernetes updates in
place; a subPath mount would never change and must not be introduced. A file that fails validation (bad slug, unknown client_type, source: own
with a non-chainstack provider_id, a chain not in the registry) is logged and ignored until it changes again.
The retired oracle.chain_ids key is ignored with a warning: expected chain ids come from the registry.
oracle.nats_url changes still need a restart; a changed oracle.connect_timeout applies to the next connect
attempt. An unknown log_level is rejected like any other invalid value.
An endpoint's scraper is picked from the first segment of its network label (corechain-mainnet → corechain,
matched against the scrapers and EVM_PROTOCOLS). A prefix none of them knows falls back to the EVM scraper
when the chain registry lists the endpoint's chain as EVM, so a chain missing from EVM_PROTOCOLS is still
scraped. Anything else is a config problem: it is logged once per config load, never scraped, and counted as
source result unsupported.
The same URL listed twice for one chain (typically under a legacy and a current label) is one source: it is
requested once per round, and both entries feed their own blockchain_network_* gauge label from that one
reading, but only the first is health-tracked and observed. Two different URLs on one host for one chain get
endpoint_ids <host>, <host>#2, …; a URL keeps its id across reloads, so removing an entry never renumbers
the others. An id a removed entry frees can go to a new URL, which starts from nothing: the removed URL's
identity, health and last observation are dropped before the new one is scraped under that id. Two entries
that set the same endpoint_id for one chain are rejected. A restart numbers inferred ids afresh.
Link to infra-helm-charts
Content type
Image
Digest
sha256:127b0d18a…
Size
96.9 MB
Last updated
20 days ago
docker pull chainstack/blockchain-network-exporter