Sign inSign up

chainstack/blockchain-network-exporter

By chainstack

•Updated 20 days ago

Image
0

9.9K

chainstack/blockchain-network-exporter repository overview

⁠Blockchain network exporter

A service which exports main metadata of many blockchain networks as Prometheus metrics, and publishes per-endpoint Chain-State Oracle observations (ADR-190).

⁠Quickstart

  • poetry install
  • cp ./config/config.example.yaml ./config/config.yaml
  • python3 -m aiohttp.web -H 0.0.0.0 -P 8000 src.main:init

Tests: poetry run pytest.

⁠Metrics

Network-level (unchanged, read by chainy's head-lag check — monotonic high-water mark across all sources):

  • blockchain_network_block_number{network} - The current block height of the blockchain
  • blockchain_network_block_timestamp{network} - The timestamp of the last block height update

Per-endpoint oracle attribution (hashes never appear in metrics; they travel over NATS only):

  • oracle_endpoint_client_info{chain,endpoint_id,source,provider_id,client_type} - always 1; client implementation and attribution of each observed endpoint
  • oracle_endpoint_client_build_info{chain,endpoint_id,source,provider_id,client_type,client_build,client_version} - always 1; the build and version each observed endpoint runs, derived from its web3_clientVersion by the contract's parse_build_info (conduit-op-reth/v2.2.3-38f6b93/… gives conduit-op-reth, 2.2.3-38f6b93), at most 64 bytes each, unknown when nothing qualifies. client_type stays the diversity signal: a vendor build is an alias of its client, so the build is visible here without counting as a second implementation. The raw string stays on NATS. Replaced on upgrade, withdrawn with client_info
  • oracle_endpoint_head_lag_blocks{chain,endpoint_id,source,provider_id} - blocks the endpoint's latest reading sits behind the highest latest reading among the chain's current sources (sources excluded from observations do not count; with a single source it is 0). Set once per scrape round, after every source has reported, so all sources of a chain are measured against the same reference. Sources are read at slightly different moments, so on fast chains it carries a few blocks of noise; alert on min_over_time(…[1m]), which a source that is really behind never brings to 0
  • oracle_endpoint_head_age_seconds{chain,endpoint_id,source,provider_id} - age of the endpoint's latest block when it was read (read time minus block timestamp). It needs no peers, so it also catches a sole source that answers but is stuck, and it is in seconds on every chain. Lag against peers is a query, free of read-time skew and of pauses in block production (a quiet chain ages every source alike): oracle_endpoint_head_age_seconds - on (chain) group_left () min by (chain) (oracle_endpoint_head_age_seconds). Block timestamps have 1s resolution; chains that only produce blocks on demand (Avalanche C-chain, quiet testnets) age while healthy, so alert on the relative form there. The absolute form trusts the exporter's clock
  • oracle_endpoint_finality_delay_seconds{chain,endpoint_id,source,provider_id,tag} - how far the endpoint's safe / finalized head sits behind its latest, from block timestamps of its last safe/finalized fetch (only while publishing). One that keeps growing is a source whose finality has stopped; across a chain's sources it is the measured input for the aggregator's per-network settle depth (CORE-15327). Withdrawn with the other head series
  • oracle_observations_total{chain,result} - outcome per scrape: published, stored (no NATS configured), publish_failed, unknown_chain (label not in the chain registry), no_hash (protocol without hash capture yet), no_chain_id, chain_id_mismatch, invalid

chain on every oracle and per-source metric is the chain-state-oracle registry slug (see below), not the network label of the legacy gauges; while a label is unresolved it falls back to the label itself.

Per-source scrape health (CORE-15586) — effective participation, not configured count:

  • blockchain_network_source_scrapes_total{chain,endpoint_id,source,provider_id,result} - success or a failure class: dns, tls, connect, timeout, network, http_429, http_4xx, http_5xx, rpc_error, parse, unsupported (no scraper for the label, see Configuration), other
  • blockchain_network_source_last_success_timestamp_seconds{…} - unix time of the last successful scrape, 0 if never
  • blockchain_network_source_up{…} - 1 while the latest scrape succeeded and is fresh, 0 when failing or stale
  • blockchain_network_source_excluded{…,reason} - 1 while the source scrapes fine but is kept out of observations: chain_id_mismatch, no_chain_id (the identity probe has not yielded a chainId yet) or invalid. Absent otherwise, and cleared as soon as a scrape fails or goes stale, so it is only ever set where up is 1. Sources of chains that emit no observations at all (not in the registry, or no hash capture yet) are never marked: they contribute to the gauges

up is scrape health; up minus excluded is what actually contributes. The or … * 0 keeps a chain with no excluded source in the result instead of dropping it from the subtraction:

sum by (chain) (blockchain_network_source_up)                          # sources scraping successfully
sum by (chain) (blockchain_network_source_up)
  - sum by (chain) (blockchain_network_source_excluded or blockchain_network_source_up * 0)
                                                                       # sources contributing right now (0, not absent, when none)
blockchain_network_source_excluded == 1                                # which sources are kept out, and why
time() - blockchain_network_source_last_success_timestamp_seconds      # seconds since each source last worked
rate(blockchain_network_source_scrapes_total{result!="success"}[5m])   # failure rate by class

A source with no success in three scrape rounds in a row (or for oracle.stale_after_seconds, when set) is dropped from /observations and loses its oracle_endpoint_client_info series, so a dead source stops looking like a quorum member; its counters and last-success timestamp remain. Rounds rather than seconds by default: a round lasts as long as its slowest endpoint, so a fixed TTL would count one slow round as several missed ones. A source that still scrapes but whose scrape yields no observation (chain_id_mismatch, no_hash, invalid, …) is withdrawn the same way at once (and, for chain_id_mismatch, no_chain_id and invalid, marked on blockchain_network_source_excluded), and a chain_id_mismatch is re-probed on identity_retry_interval so a source fixed upstream returns within a minute. The legacy blockchain_network_* gauges keep their high-water-mark semantics until chainy moves off them (CORE-15337).

⁠Oracle observations

An EVM probe sends the single eth_getBlockByNumber(latest) call this exporter has always made, and the latest block's hash comes with it. Only where observations are published (oracle.nats_url set) does it also fetch safe and finalized, batched with latest. An exporter that does not publish sends exactly what it sent before the oracle.

How often is measured, not configured per chain. Each fetch measures the chain's finality delay: how far its safe or finalized head sits behind latest (whichever is closer), from the block timestamps of that one batch, and the smallest across the chain's sources, so one stuck source cannot slow its chain down. The next fetch comes a quarter of that delay later, which keeps the reported heads within a quarter of where the chain's own finality already puts them, bounded by oracle.finality_interval_min (60s) and oracle.finality_interval_max (10 min). Measured on 2026-09-24 against public endpoints of the chains in the prod config: instant-finality chains (BSC, Polygon, Avalanche, Sonic) sit at the 60s floor, OP-stack L2s at 60s–2 min (their safe moves every few minutes), Ethereum and Arbitrum at about 2 min, and zk rollups (Linea, zkSync, Scroll) at the 10 min ceiling. A source that answers neither tag waits the ceiling instead of paying for a useless batch every minute. Each source's first fetch falls at a random point within the floor, so the batches spread over the rounds.

A batch costs two calls more than the single call. A provider that rejects it gets the single call instead, and only that call for an hour. Rejecting means a 4xx other than 429, a reply that is not a batch answer, or three batches in a row that came back with nothing but errors while the single call worked. A batch that fails (429, 5xx, timeout) never fails the scrape: the batch exists for the oracle only, so the gauges hang on the single call, which the provider then gets alone for ten minutes, the pre-oracle load on a provider that is rate-limiting or down, while safe/finalized come from the chain's other sources.

Each probe emits one observation per endpoint — {number, hash} for latest, and for safe/finalized on the scrapes that fetched them (between fetches they are absent, never repeated stale), tagged source, providerId, clientType, chainId — encoded with the canonical JSON profile of chain-state-oracle⁠ (schemaVersion 1.1.0).

  • Published to NATS subject oracle.<chain>.obs when oracle.nats_url is set. Publishing is fail-open: a NATS outage is counted and logged, and never affects the gauges. The connection is made in the background (oracle.connect_timeout per attempt, retried every 10s), so an unreachable NATS never delays a scrape; the warm-up scrape waits for the first attempt to settle, one bounded attempt at most, so a healthy deploy does not count its first round as publish_failed.
  • Always served on GET /observations[?chain=<slug or alias>] as a JSON array of the last payload per endpoint, byte-identical to what goes on the wire (?network= is accepted as a synonym).
  • Observations are reported, not max-merged: a lower reading from a source is a valid observation.
  • Failed scrapes are never encoded as observations; they are counted in blockchain_network_source_scrapes_total.
  • Client identity (web3_clientVersion, eth_chainId) is probed once per identity_probe_interval, not per scrape, because aggregators rate-limit extra calls; each source's first re-probe falls at a random point in the second half of the interval, so a fleet probed in one warm-up round does not re-probe in one round every hour. A method the provider refuses (4xx) is logged with its status and left out of the identity; without eth_chainId the source is excluded as no_chain_id and re-probed on identity_retry_interval. Normalization mirrors the contract's client-identity table (schema 1.1.0 covers the served fleet: nitro, sonic, kaia, avalanchego, …). Proxy rewriters normalize to unknown. client_type in config overrides the probe (read on every scrape, so a reload applies it at once, without waiting for the next probe); it is needed for reth-derived distributions that self-report reth (base-reth, nanoreth, Plasma) and for older op-reth builds. Aggregators such as dRPC answer with a synthetic string (Geth/v10.0.0/drpc) that would classify as geth; pin them to client_type: unknown so they never count toward client diversity.
  • A source whose eth_chainId differs from the registry's pin for its chain is excluded from observations.
  • Non-EVM scrapers do not capture hashes yet and emit no observations (Solana: CORE-15585).

Conformance against the shared golden vectors is tested in tests/test_observation.py; see tests/vectors/SOURCE.md to refresh them.

⁠Chain registry

The chain an observation carries, and the oracle.<chain>.obs subject it goes to, is a slug from the chain-state-oracle chain registry (src/oracle/chains.json, generated upstream from the platform catalog, chainstack/core networks.csv): one slug per Chainstack internal network name with the chain id (or Solana genesis hash) it pins. Producers never invent a subject, so own and external observations for one chain always meet on one subject even when their Prometheus labels differ (polygon-mainnet is a registry alias of polygon-pos-mainnet; kaia-mainnet of klaytn-mainnet).

Each endpoint's chain is resolved at config load, in this order: an explicit chain: (must be a registry slug or alias, otherwise the config is rejected), then additional_network_label, then network. A label the registry does not know is logged at load, keeps feeding the blockchain_network_* gauges, and is counted as unknown_chain instead of emitting an observation. The fix is a catalog row in chainstack/core followed by a registry refresh upstream, or chain: pointing at the slug the label means. The legacy gauges keep the network label unchanged either way.

src/oracle/registry.py mirrors the Go and Rust lookup semantics and passes the same shared fixture (tests/vectors/registry/resolve_cases.json, tests/test_registry.py).

⁠Configuration

oracle:
  nats_url: "nats://jetstream-nats:4222"   # omit to disable publishing
  finality_interval_min: 60                # bounds of the measured safe/finalized fetch interval, seconds;
  finality_interval_max: 600               # only while publishing
endpoints:
  - network: ethereum-mainnet
    url: https://lb.drpc.org/ogrpc?network=ethereum&dkey=...
    chain: ethereum-mainnet   # registry slug or alias; resolved from additional_network_label / network when omitted
    source: external          # inferred from host when omitted (*.chainstack.com, *.p2pify.com, *.chainstacklabs.com → own)
    provider_id: drpc         # inferred from the host's registrable name when omitted (rpc.example.co.uk → example);
                              # IP and cluster-internal hosts infer `unknown` (set it)
    endpoint_id: lb.drpc.org  # defaults to the host; unique per chain when set; never put credentials here, it goes on the wire
    client_type: op-reth      # optional override of the probed client; must be a clientType enum value

The config file is re-read when its mtime changes, so a provider can be added or dropped without an image deploy. The Helm chart mounts the config Secret as a directory (not a subPath), which Kubernetes updates in place; a subPath mount would never change and must not be introduced. A file that fails validation (bad slug, unknown client_type, source: own with a non-chainstack provider_id, a chain not in the registry) is logged and ignored until it changes again. The retired oracle.chain_ids key is ignored with a warning: expected chain ids come from the registry. oracle.nats_url changes still need a restart; a changed oracle.connect_timeout applies to the next connect attempt. An unknown log_level is rejected like any other invalid value.

An endpoint's scraper is picked from the first segment of its network label (corechain-mainnet → corechain, matched against the scrapers and EVM_PROTOCOLS). A prefix none of them knows falls back to the EVM scraper when the chain registry lists the endpoint's chain as EVM, so a chain missing from EVM_PROTOCOLS is still scraped. Anything else is a config problem: it is logged once per config load, never scraped, and counted as source result unsupported.

The same URL listed twice for one chain (typically under a legacy and a current label) is one source: it is requested once per round, and both entries feed their own blockchain_network_* gauge label from that one reading, but only the first is health-tracked and observed. Two different URLs on one host for one chain get endpoint_ids <host>, <host>#2, …; a URL keeps its id across reloads, so removing an entry never renumbers the others. An id a removed entry frees can go to a new URL, which starts from nothing: the removed URL's identity, health and last observation are dropped before the new one is scraped under that id. Two entries that set the same endpoint_id for one chain are rejected. A restart numbers inferred ids afresh.

⁠Helm chart

Link to infra-helm-charts⁠

Tag summary

Content type

Image

Digest

sha256:127b0d18a…

Size

96.9 MB

Last updated

20 days ago

docker pull chainstack/blockchain-network-exporter