When to Use Micromegas¶
Micromegas is an open-source (Apache-2.0), self-hosted observability stack that collects logs, metrics and traces from native and client processes, stores raw payloads in object storage with PostgreSQL metadata, and is queried with SQL over Apache Arrow FlightSQL.
TL;DR. Choose Micromegas for low instrumentation overhead and low cost, and for the high-frequency, full-resolution telemetry that efficiency makes affordable. Choose a peer when you want the established, widely adopted default, or when you do not want to operate the stack yourself. The summary at the end lists when each choice fits.
Last reviewed: October 2026
Every claim on this page links to its source: the project's own docs or repository for the peers, and the Micromegas docs or source code for Micromegas. If something is wrong or out of date, please open an issue.
Micromegas in brief¶
- Instrumentation. Native SDKs for Rust (the
micromegas-tracingmacros such asspan_scope!, inmacros.rs, and#[span_fn], inlib.rs) and an Unreal Engine plugin record spans, logs and metrics. A C ABI records logs and metrics and makes Micromegas easy to embed in any app or language that can load a shared library (lib.rs); the Blender add-on is an example, calling it from Blender's embedded Python throughctypes. Thetracing-crate interop captures existingtracingevents as logs (tracing_interop.rs). Events are recorded in-process on the calling thread, the telemetry sink batches and ships them off the hot path, and the intent is instrumentation that stays on in production. - Ingestion. Ingestion is HTTP: the native transit format and OTLP/HTTP (protobuf or JSON, gzip) for logs, metrics and traces (ingestion service). The SDKs batch events per stream and send them as LZ4-compressed blocks (
InsertBlockRequest.cppin Unreal); the Rust sink flushes a stream when its buffer fills (by default 10 MiB for logs and thread spans, 1 MiB for metrics) or every 60 s (lib.rs,flush_monitor.rs,stream_block.rs). The ingestion service stores each block as received: one object-storage write and one PostgreSQL row cover a whole block of up to hundreds of thousands of events, with no per-event parsing (web_ingestion_service.rs). Events are parsed afterwards: logs and metrics continuously, spans when queried. Queries are served over FlightSQL (FlightSQL server). - Analytics. Raw payloads stay in object storage (S3, GCS or local) and metadata lives in PostgreSQL (architecture). A lakehouse materializes views to Parquet and queries them with Apache DataFusion (lakehouse architecture, query guide). Logs and metrics are materialized continuously into global views by the maintenance daemon; spans and per-process views are materialized only when queried (JIT ETL, on-demand processing), so their processing cost follows what is queried.
- Derived views. An admin defines a new materialized view at runtime with
CREATE MATERIALIZED VIEW: SQL over the lakehouse tables, stored in PostgreSQL and picked up by the running services without a restart (materialized views). Every second, the maintenance daemon materializes each view from newly ingested data, then merges those partitions into minute, hour and day partitions (maintenance.rs). A reduced view, such as per-minute log counts by process and level, lets a fleet-wide dashboard over weeks read a handful of small partitions instead of scanning terabytes of raw events, and drilling down reads full-resolution rows only for the processes and time range selected (JIT ETL). Reduced views sit beside the raw data instead of replacing it, so old data is not downsampled: every event keeps full resolution until the retention horizon, 90 days by default (maintenance). When late data changes a partition's source, the partition is rebuilt from the raw data. View definitions can be kept in git with aplan/applyworkflow (views as code), and the built-inlog_statsview is defined this way. - Presentation. The analytics web app runs notebooks. Its server fetches query results over FlightSQL and streams them to the browser as Arrow IPC record batches (
stream_query.rs,arrow-stream.ts). The browser keeps them in Arrow format in memory, where DataFusion, compiled to WASM, queries them locally, so later cells can query earlier cells' results without going back to the server (execution model). A Grafana data source plugin covers dashboards, and Grafana is also the path for alerting. A process's spans can be exported as a Perfetto trace (perfetto_trace_chunks). - Access control. Every row is stamped server-side with an audience taken from the ingestion credential, so a producer cannot forge it (audience stamping). Read grants are separate and editable, and re-sharing applies to already-ingested data within the grant-cache TTL (60 s by default), with no restamping (authorization). All of it is in the Apache-2.0 build, with no paid tier.
Limits, stated plainly:
- PostgreSQL is required for metadata (architecture).
- OTLP is HTTP-only; there is no OTLP/gRPC (OTLP limitations).
- There is no built-in alert engine; alerting goes through Grafana (Grafana plugin).
- Among the native SDKs, spans come from Rust and Unreal; the C ABI records logs and metrics only. OTLP/HTTP traces are also accepted as spans (native SDK, Unreal plugin).
- There is no browser or mobile-web RUM SDK and no session replay; RUM for game clients comes through the Unreal plugin.
- Micromegas is a younger project: it was extracted in January 2024 from the Legion Labs engine, where its telemetry code was developed from 2021 to 2022. Its community is smaller than its peers'; issues and design discussions go straight to the maintainers.
- You operate it yourself: PostgreSQL, object storage, and the services (or the single-process monolith).
At a glance¶
Each project's section below gives the sources for its row.
Deployment and licensing
| Project | License (OSS edition) | Storage | Also requires |
|---|---|---|---|
| Micromegas | Apache-2.0, no paid tier | Raw payloads in object storage; Parquet views | PostgreSQL; Grafana with the Micromegas plugin for alerts |
| Parseable | AGPL-3.0; paid PromQL, HA | Parquet on object storage | Nothing |
| OpenObserve | AGPL-3.0; paid SSO, advanced RBAC | Parquet on object storage | Nothing on one node; PostgreSQL and NATS for HA |
| GreptimeDB | Apache-2.0; paid alerting, RBAC | Parquet on object storage | Nothing standalone; etcd, PostgreSQL or MySQL when distributed |
| SigNoz | MIT with proprietary ee/; AGPL collector |
ClickHouse | ClickHouse Keeper or ZooKeeper; SQLite |
| ClickHouse / ClickStack | Apache-2.0; HyperDX MIT | Local disk, S3 as tiered disk | Keeper for replication; MongoDB for ClickStack |
| Grafana LGTM | AGPL-3.0; paid GEL, GET, GEM | Object storage | Kafka for Mimir's Helm default and Tempo microservices |
| InfluxDB 3 Core | MIT or Apache-2.0; paid HA | Parquet on object storage | Nothing |
| VictoriaMetrics | Apache-2.0; paid downsampling, mTLS | Local disk | Nothing |
| Quickwit | Apache-2.0 | Index splits on object storage | PostgreSQL when distributed |
| Prometheus (Thanos) | Apache-2.0 | Local TSDB; Thanos adds object storage | Alertmanager for alerts |
Capabilities
| Project | Signals | Query | Own in-process SDKs | UI / alerting |
|---|---|---|---|---|
| Micromegas | Logs, metrics, traces | SQL | Rust, Unreal Engine, C ABI | Notebooks; alerts via Grafana |
| Parseable | Logs, metrics, traces, events | SQL; PromQL paid | None (OTel); small Go SDK | Built-in; threshold alerts |
| OpenObserve | Logs, metrics, traces, RUM, session replay, profiles | SQL, PromQL, full-text | RUM SDKs; OTel for backends | Built-in; alerts |
| GreptimeDB | Metrics, logs, traces | SQL, PromQL | Ingest clients only | Dashboard; alerting paid |
| SigNoz | Logs, metrics, traces, exceptions | Query builder, PromQL, ClickHouse SQL | None (OTel) | Built-in APM; alerts |
| ClickHouse / ClickStack | Logs, metrics, traces | ClickHouse SQL; HyperDX search | None (OTel) | HyperDX; alerts |
| Grafana LGTM | Logs, traces, metrics, profiles | LogQL, TraceQL, PromQL | None (OTel); Faro, Beyla, Pyroscope | Grafana; Grafana Alerting |
| InfluxDB 3 Core | Metrics, events; logs and traces via Telegraf | SQL, InfluxQL | Write clients only | Explorer; plugin alerts |
| VictoriaMetrics | Metrics, logs, traces | MetricsQL, LogsQL | Go metrics package |
vmui; vmalert |
| Quickwit | Logs, traces | Elasticsearch-compatible API | None (OTel) | Basic UI; Grafana, Jaeger |
| Prometheus (Thanos) | Metrics | PromQL | Metric client libraries | Expression browser; Alertmanager |
Micromegas vs. Parseable¶
Parseable is a Rust "unified observability platform on a data lake architecture" for logs, metrics, traces and events (repo). It is AGPL-3.0; paid Cloud and Enterprise tiers gate PromQL, the HA cluster, APM, anomaly detection and AI features, while SQL, dashboards, threshold alerts, OIDC/SSO and RBAC are in the open-source edition (pricing).
Choose Parseable when you want one binary with no metadata database, with all signals stored as Parquet on object storage and queried with SQL, and broad ingestion compatibility: its own HTTP JSON API, OTLP over HTTP, Kafka, Fluent Bit, Vector, Logstash and Filebeat (integrations, architecture).
How Micromegas differs. Micromegas and Parseable both run SQL on DataFusion over Parquet (Parseable Cargo.toml), so that axis is shared. Parseable relies on OTel for instrumentation, with a small Go SDK; Micromegas ships in-process SDKs for Rust and Unreal Engine. The ingestion paths differ as well. Parseable parses each request into JSON values, infers and merges a schema over every record and converts the batch to Arrow, staged on local disk and turned into Parquet each minute (json.rs, architecture). Its open-source edition accepts OTLP as JSON only, so OTel exporters and collectors must be set to send JSON (ingest_utils.rs, OTLP logs). Micromegas's SDKs send batched, LZ4-compressed binary blocks. Per request, its ingestion service does a constant amount of work for a block of up to hundreds of thousands of events (one object write and one PostgreSQL row), where Parseable parses, schema-checks and converts every event (ingestion). Logs and metrics are parsed afterwards, continuously; spans and per-process views skip parsing until someone queries them, so for most span data that work is never done. Micromegas stores every emission as its own row. In the open-source edition Parseable's distributed mode allows many ingest nodes but only one query node (OSS Helm). Micromegas needs PostgreSQL for metadata; Parseable does not.
Micromegas vs. OpenObserve¶
OpenObserve is a Rust backend with a Vue UI covering logs, metrics, traces, RUM, session replay, profiles and LLM observability (repo). It is AGPL-3.0 in the open-source edition (it moved from Apache); the Enterprise edition is under a commercial license, free up to 50 GB/day (license) and gates SSO, advanced RBAC, audit logs, federation and AI features (features). It has an LLM observability feature set for agent traces.
Choose OpenObserve when you need the broadest signal coverage (browser and mobile RUM and session replay included), a rich built-in UI with dashboards, pipelines, alerts and incidents, full-text search via Tantivy, or an easy migration off ELK through its Elasticsearch-compatible _bulk API (ingestion, metrics).
How Micromegas differs. Both run SQL on DataFusion over Parquet on object storage (OpenObserve Cargo.toml). OpenObserve uses OTel SDKs for backend code, plus its own RUM SDKs (repo); Micromegas has in-process Rust and Unreal SDKs. On ingestion, OpenObserve turns each event into a JSON value, flattens it and resolves the batch schema before converting it to Arrow (ingest.rs, mod.rs); the resulting Parquet is merged again on upload, where a Tantivy index is built, and compacted later (parquet.rs, config.rs). Micromegas stores each compressed block of up to hundreds of thousands of events as received, with no per-event work (ingestion). OpenObserve writes Parquet at ingestion, whereas Micromegas processes spans and per-process views only when queried. OpenObserve keeps metadata in SQLite on one node, and PostgreSQL plus NATS in HA mode (architecture). Micromegas also offers one SQL surface rather than SQL plus PromQL, and notebooks running the same engine in the browser (WASM). OpenObserve's Enterprise edition gates advanced RBAC (features), whereas Micromegas's per-row access control is in the open-source build. Where OpenObserve offers profiling, Micromegas's difference is spans named by the data being processed rather than sampled call stacks (code vs. data).
Micromegas vs. GreptimeDB¶
GreptimeDB is a Rust "observability database": one columnar engine for metrics, logs and traces, with SQL joins across signals (repo). The core is Apache-2.0 under an open-core model; Enterprise gates Triggers (alerting), LDAP, RBAC, audit logs and automatic rebalancing (enterprise, triggers).
Choose GreptimeDB when you want a long-term replacement for Prometheus storage, want to migrate one signal at a time (it ingests OTLP, Prometheus remote write, Loki push, Elasticsearch _bulk, InfluxDB line protocol and the MySQL and PostgreSQL wire protocols (docs)), or have high-cardinality metrics. Standalone mode is a single binary (README).
How Micromegas differs. Both use SQL on DataFusion over Parquet in object storage (config). GreptimeDB ships ingester client libraries only, with no instrumentation SDKs; Micromegas has in-process SDKs. On ingestion, GreptimeDB's OTLP, line-protocol and JSON paths build a row of values per event, and JSON logs go through a dynamic value tree and schema resolution per event (logs.rs, greptime.rs); its fastest path, bulk Arrow over Flight, needs a client that builds Arrow batches (protocol benchmark). Micromegas stores each compressed block of up to hundreds of thousands of events as received, with no per-event work (ingestion). GreptimeDB writes Parquet at ingestion; Micromegas processes spans and per-process views only when queried. Micromegas also offers notebooks running the same engine in the browser (WASM). GreptimeDB's Enterprise edition gates RBAC (enterprise), whereas Micromegas's per-row access control is in the open-source build. Distributed GreptimeDB needs etcd, PostgreSQL or MySQL for metasrv, plus optional Kafka for the WAL; Micromegas needs PostgreSQL.
Micromegas vs. SigNoz¶
SigNoz is a Go and React OTel-native APM covering logs, metrics, traces, exceptions and LLM observability (repo). The code is MIT outside ee/ and cmd/enterprise/, which are proprietary, and its collector is AGPL-3.0 (LICENSE, collector). Cloud and Enterprise gate anomaly detection, SAML, fine-grained RBAC and audit logs (pricing).
Choose SigNoz when you want the best out-of-the-box APM experience for OTel-instrumented services, with built-in APM views, traces, logs, dashboards and alerts and no Grafana needed (architecture). It also has LLM observability features for agent workloads.
How Micromegas differs. SigNoz relies on OTel SDKs only; Micromegas adds in-process SDKs for Rust and Unreal Engine. SigNoz stores data in ClickHouse, with ClickHouse Keeper or ZooKeeper and SQLite for dashboards, alerts and users (moldings); Micromegas stores Parquet on object storage with PostgreSQL metadata. On ingestion, SigNoz always goes through its collector, which decodes OTLP, rebuilds each log record's attributes as maps and serializes them to JSON only to meter their size, then issues five inserts per log batch (exporter.go); ClickHouse then builds token and n-gram bloom filters on each insert (logs schema). Micromegas stores each compressed block of up to hundreds of thousands of events as received, with no per-event work (ingestion). Micromegas keeps raw payloads in object storage and processes spans and per-process views only when queried. Every emission is stored as its own row. SigNoz offers a query builder, PromQL and ClickHouse SQL (architecture); Micromegas has one SQL surface, plus notebooks running the same engine in the browser (WASM). SigNoz Cloud and Enterprise gate fine-grained RBAC (pricing), whereas Micromegas's per-row access control is in the open-source build.
Micromegas vs. ClickHouse and ClickStack¶
ClickHouse is a C++ columnar OLAP database; its own docs say it "isn't an out-of-the-box solution for Observability" but is a highly efficient storage engine (intro). ClickStack bundles ClickHouse, the HyperDX UI and an OTel collector distribution (overview). ClickHouse is Apache-2.0 and HyperDX is MIT (ClickHouse, HyperDX).
Choose ClickHouse when you need raw query speed and compression at very large scale and are prepared to build your own pipeline, with the benefit of a mature ecosystem; ClickStack's HyperDX adds built-in search, traces, dashboards and alerts. ClickStack accepts OTLP over HTTP and gRPC through its collector (docs).
How Micromegas differs. ClickHouse stores MergeTree data on local disk, with S3 as a tiered disk (S3); replication needs ClickHouse Keeper or ZooKeeper, and self-hosted ClickStack also needs MongoDB for dashboards, saved searches and alerts (deployment). Micromegas keeps raw payloads in object storage with PostgreSQL metadata, and processes spans and per-process views only when queried. ClickStack's instrumentation is OTel-based; Micromegas adds in-process SDKs. On ingestion, ClickStack's collector runs transforms on every log, parsing JSON bodies and inferring severity with regular expressions (config.yaml), and ClickHouse sorts, indexes and compresses each insert into a part that background merges rewrite later (insert strategy). ClickHouse measured the collector hop on its own internal logs: OTel collectors used "over 800 CPU cores to ship 2 million logs per second", against 70 cores for 37 million with a byte-for-byte copy (LogHouse). Micromegas stores each compressed block of up to hundreds of thousands of events as received, with no per-event work (ingestion). Micromegas stores every emission as its own row, with one SQL surface and notebooks running the same engine in the browser (WASM). ClickHouse's open-source build has row policies for row filtering; the Micromegas difference is that each row's audience is stamped server-side from the ingestion credential rather than set by the producer (audience stamping). ClickHouse's incremental materialized views are "a trigger that runs a query on blocks of data as they're inserted" (materialized views), and backfilling one means re-inserting the source data; Micromegas derived views are materialized per time partition from the raw data it keeps, so late data and redefinitions are rebuilt the same way.
Micromegas vs. Grafana LGTM (Loki, Tempo, Mimir, Pyroscope)¶
Grafana LGTM is one backend per signal viewed in Grafana: Loki indexes labels, not log contents; Tempo stores traces; Mimir is long-term Prometheus storage; Pyroscope adds continuous profiling. It is AGPL-3.0, with Apache-2.0 exceptions in LICENSING.md (Loki); paid GEL, GET and GEM add tenant management, token auth and cross-tenant query (GEL, GET, GEM).
Choose Grafana LGTM when you want the de-facto self-hosted standard: object-storage backends (Loki storage), purpose-built query languages including PromQL (LogQL, TraceQL), a mature UI and alerting (Grafana Alerting, plus the Mimir ruler and Alertmanager), and an OTel-first approach via Alloy. Each backend runs as one binary with -target=all or as microservices (Loki, Mimir).
How Micromegas differs. Micromegas is one system where LGTM is several. LGTM runs one backend per signal, each deployed, scaled and configured with its own storage, and each queried in its own language (LogQL, TraceQL, PromQL), with no SQL. Grafana ties the signals together by configuring links between data sources: trace to logs maps span attributes to Loki labels and generates a LogQL query, and the way back relies on applications writing trace IDs into their log lines (trace to logs). Micromegas has one ingestion service, one store and one SQL engine for logs, metrics and traces (architecture). Every signal is keyed by the same process and stream identifiers, so one query can join logs, metrics and spans (schema reference), and one access-control model covers all of them (authorization).
LGTM relies on upstream OTel SDKs, Faro, Beyla and Pyroscope profiling SDKs (otel docs); Micromegas has in-process SDKs for Rust and Unreal Engine. On ingestion, Loki and Mimir parse each stream or series, hash it to a shard and write it to three ingesters by default, each with its own WAL (Loki, Mimir); in microservices mode Tempo regroups spans by trace and writes them through a Kafka-compatible queue (architecture). Micromegas stores each compressed block of up to hundreds of thousands of events as received, with no per-event work (ingestion). The mimir-distributed Helm chart enables Kafka-based ingest storage by default, which "requires a production-grade Apache Kafka cluster" (ingest storage), and Tempo's microservices mode requires a Kafka-compatible system (modes); Micromegas needs PostgreSQL and object storage. Micromegas stores every emission as its own row and offers notebooks running the same engine in the browser (WASM). Access control in LGTM is a paid add-on per backend: Enterprise Logs adds label-based access control and Enterprise Metrics adds fine-grained access control (GEL, GEM); Micromegas's per-row access control is in the open-source build. Against Pyroscope's sampled call stacks, Micromegas spans can be named by the data being processed (code vs. data). Micromegas also ships a Grafana data source plugin, so the two can be combined.
Micromegas vs. InfluxDB 3 Core¶
InfluxDB 3 Core is a time-series database built for recent data, with last-value and distinct-value caches and an embedded Python processing engine (docs). It handles metrics and events first; logs and traces arrive only via Telegraf conversion. It is MIT or Apache-2.0 (repo); commercial Enterprise adds HA, read replicas, multi-node, long-range historical queries and historical compaction (product). Core queries cover about 72 hours by default (query-file-limit), raisable at a memory and speed cost (query, config).
Choose InfluxDB 3 Core when you want a permissive license, the same Arrow, DataFusion, Parquet and Flight SQL stack, all metadata in object storage (setup, backup), very fast recent-data queries, and an embedded Python engine. It queries with SQL and InfluxQL over HTTP, Arrow Flight and Flight SQL (query).
How Micromegas differs. InfluxDB 3 Core and Micromegas both run DataFusion over Parquet, and Micromegas owes InfluxData thanks for it: InfluxData has invested heavily in Apache DataFusion and the Rust Arrow implementation, contributing upstream and often leading the work (FDAP architecture, InfluxDB 3 GA), and Micromegas's query engine builds on that work. InfluxDB 3 ingests line protocol only, with OTLP converted by Telegraf (write, Telegraf OTel), and its client libraries are for writing and querying, not instrumentation; Micromegas accepts OTLP/HTTP natively and has in-process SDKs. On ingestion, InfluxDB 3 Core parses line-protocol text and resolves every tag and field against its catalog, one line at a time (validator.rs), and it groups all writes into one WAL file in object storage per second (object_store.rs). Micromegas stores each compressed block of up to hundreds of thousands of events as received, with no per-event work (ingestion). Core is a single node with retention fixed per database at creation (retention); Micromegas runs independently scaled services over object storage with PostgreSQL metadata. Micromegas processes spans and per-process views into Parquet only when queried, and also offers notebooks running the same engine in the browser (WASM).
Micromegas vs. VictoriaMetrics, VictoriaLogs and VictoriaTraces¶
VictoriaMetrics is a family of three Go databases from one vendor: VictoriaMetrics (Prometheus long-term storage), VictoriaLogs, and VictoriaTraces, which is built on VictoriaLogs and stores spans as structured logs (VictoriaTraces docs). All three are Apache-2.0, cluster versions included (cluster); Enterprise gates downsampling, multiple retentions, backup automation, mTLS and anomaly detection (enterprise). VictoriaTraces is pre-1.0 and warns that its APIs "may not be backward compatible" (repo).
Choose VictoriaMetrics when you want drop-in compatibility with Prometheus, Loki, Elasticsearch and Jaeger clients, operational simplicity ("a single small executable without external dependencies"), an open-source cluster mode, and low resource use. It ingests Prometheus remote write and scraping, Influx line protocol, OTLP over HTTP and more (VictoriaMetrics, VictoriaLogs), and vmalert evaluates MetricsQL and LogsQL rules (vmalert).
How Micromegas differs. VictoriaMetrics stores data on local disk and uses object storage for vmbackup snapshots only (vmbackup); it does not use Parquet. Micromegas keeps raw payloads in object storage and processes spans and per-process views into Parquet only when queried. VictoriaMetrics queries with MetricsQL ("backwards-compatible with PromQL", metricsql) and LogsQL, with no SQL (faq); Micromegas has one SQL surface and notebooks running the same engine in the browser (WASM). VictoriaMetrics relies on OTel and Prometheus clients, with one first-party Go metrics package; Micromegas has in-process SDKs for Rust and Unreal Engine. VictoriaMetrics has a lean ingest path: a cached series lookup per row, no WAL, and data buffered in memory for a few seconds, at the cost that the last few seconds may be lost on an unclean shutdown (docs). Micromegas stores each compressed block of up to hundreds of thousands of events as received, with no per-event work (ingestion). Micromegas stores every emission as its own row. VictoriaMetrics needs no PostgreSQL; Micromegas does.
Micromegas vs. Quickwit¶
Quickwit is a Rust search engine built on Tantivy for logs and traces, with compute separated from storage and search running directly on object storage (overview). It is Apache-2.0 (relicensed from AGPL when Datadog acquired the team in January 2025); the founders said they would focus on "building a new product with Datadog", and there is no standalone commercial offering or paid support (announcement). It is still being released: v0.9.1 shipped on 2026-09-23 (releases).
Choose Quickwit when you want fast full-text log search on cheap object storage, a drop-in for Elasticsearch-compatible tooling, or Jaeger-native trace storage. It ingests OTLP for logs and traces, Jaeger, an Elasticsearch-compatible API, Kafka and SQS (v0.9.0 notes).
How Micromegas differs. Quickwit and Micromegas both use object storage and can keep metadata in PostgreSQL (metastore); that axis is shared. Quickwit targets logs and traces (docs); Micromegas stores logs, metrics and traces as Parquet. Quickwit exposes an Elasticsearch-compatible query API with no SQL; Micromegas queries with SQL. Quickwit relies on OTel SDKs; Micromegas has in-process SDKs. On ingestion, Quickwit decodes OTLP protobuf, re-encodes each record as JSON for its write-ahead log, then parses that JSON again and indexes it with Tantivy (logs.rs, doc_processor.rs); merges later rewrite splits until they reach 10 million documents (merge policy). Micromegas stores each compressed block of up to hundreds of thousands of events as received, with no per-event work (ingestion). Micromegas offers notebooks running the same engine in the browser (WASM), and a Grafana plugin.
Micromegas vs. Prometheus¶
Prometheus is "an open-source systems monitoring and alerting toolkit" (overview). It is metrics only: each sample is a float64 or native histogram with a millisecond timestamp (data model). It is Apache-2.0, written in Go, and CNCF graduated; Thanos is Apache-2.0 and CNCF incubating (Prometheus, Thanos).
Choose Prometheus when you want the de-facto standard for service metrics: PromQL, simple single-binary operation, service discovery, mature alerting through Alertmanager (alerting), and the largest exporter ecosystem, with metric client libraries for Go, Java/Scala, Node.js, Python, Ruby and Rust (client libraries). Thanos and remote storage add long-term, global views.
How Micromegas differs. Micromegas is a Prometheus alternative for sub-second, high-frequency metrics, with five differences, each set against Prometheus's own docs.
Frequency. Prometheus records one sample per series per scrape, and scrape_interval defaults to 1m (config); gauges are "snapshots of state" (instrumentation), and "If you need 100% accuracy, such as for per-request billing, Prometheus is not a good choice, as the collected data will likely not be detailed and complete enough" (overview). Micromegas stores every metric emission as its own row with a nanosecond timestamp (measures). Its built-in system monitor samples host-wide CPU usage and used and free memory every 200 ms in each process using the Rust telemetry sink or the C ABI, plus the process's own memory every 5 s (system_monitor.rs).
Dimensionality. In Prometheus "every unique combination of key-value label pairs represents a new time series… Do not use labels to store dimensions with high cardinality" (naming); "Each labelset is an additional time series that has RAM, CPU, disk, and network costs", and over 100 values it suggests "moving the analysis away from monitoring and to a general-purpose processing system" (instrumentation). Labels are therefore low-cardinality and chosen at instrumentation time. In Micromegas, process, executable, computer and user are columns on each row alongside properties, and SQL can group or filter by any of them at query time (measures). Micromegas has its own producer-side cardinality contract: metric names, log targets and property sets are interned in process memory, so they must stay bounded, and free-form values belong in the log message body (native SDK, Blender add-on).
Scale. "Prometheus's local storage is limited to a single node's scalability and durability" and "is not clustered or replicated" (storage); high availability means running "identical Prometheus servers on two or more separate machines" (FAQ). Thanos adds object storage for blocks, a global query view, deduplication and downsampling (Thanos). Micromegas runs independently scaled services over object storage (architecture).
Ingestion. Prometheus client libraries aggregate in process, so an increment is one atomic add (counter.go), and the server parses each sample on every scrape and stores it compactly, about 1–2 bytes per sample (storage); only the state at scrape time is kept. Micromegas records every emission in-process and ships it in LZ4-compressed blocks that the server stores as received (ingestion).
Precomputation. Recording rules "precompute frequently needed or computationally expensive expressions and save their result as a new set of time series", evaluated at a regular interval (recording rules). Micromegas derived views are SQL over logs, metrics and process metadata, refreshed every second, and their output is a table that can carry high-cardinality columns such as process and user. Thanos downsamples blocks older than 40 hours to 5-minute resolution and blocks older than 10 days to 1-hour resolution, with a separate retention per resolution, and warns that keeping raw data for less time than the downsampled data means "not being able to 'zoom in' to your historical data" (compactor). Micromegas keeps raw events for the whole retention period, so a long-range query reads a reduced view and can still zoom in to full resolution.
Complementary tools¶
Tracy, Unreal Insights and Perfetto give a deep view of one session; they are not head-to-head alternatives to a fleet-wide store.
- Tracy is BSD-3-Clause, with a "hybrid frame and sampling profiler" (repo).
- Unreal Insights ships with Unreal Engine under the Epic EULA (source-available, not open source) (docs).
- Perfetto is Apache-2.0, runs SQL over single trace files through trace_processor, and offers call-stack sampling (repo).
Micromegas keeps the history of many processes in a single store and makes it queryable. It also exports a process's spans as a Perfetto trace that opens in the Perfetto UI (perfetto_trace_chunks).
Code vs. data. A sampling profiler sees the call stack, so it shows which code is hot. Instrumentation can also record which data that code was processing: spans named by the asset or script (an FName in Unreal, a statically allocated string in Rust; see the Unreal instrumentation API and macros.rs), and context such as the current level, asset or route attached as an interned property set that every log and metric event references for the cost of one pointer (Default Context API, property_set.rs). In a game it is rarely the animation code that is slow; it is a particular animation. The same holds for any interpreter or resolver (script VMs, query engines, template renderers, rule engines, asset loaders, routers, dependency resolvers): the stack is the same for every input, and the cost depends on the input. This is the point against sampling profilers such as Tracy's sampler and Perfetto's call-stack sampling, and continuous profilers like Pyroscope. Unreal Insights' instrumented scopes can be named by data too; the difference there is the global view: Insights opens one process's trace at a time, while Micromegas answers a single query across every process in the fleet.
Also considered¶
- Elasticsearch / OpenSearch: Elasticsearch is AGPL, SSPL or ELv2 licensed (repo) and OpenSearch is Apache-2.0 (repo); Quickwit and OpenObserve, above, cover the Elasticsearch-compatible use case.
- Uptrace: AGPL-3.0, built on ClickHouse with PostgreSQL metadata (repo); its latest release is a beta (v2.1.0-beta.8).
- Jaeger (and Zipkin): Apache-2.0, traces only, with storage delegated to other backends (Jaeger).
- Apache SkyWalking: Apache-2.0, agent-centric APM with BanyanDB storage and GraphQL, PromQL, LogQL and TraceQL APIs (storage docs).
- Apache Doris / StarRocks: Apache-2.0 SQL warehouses, the same build-your-own category as plain ClickHouse (Doris, StarRocks).
- Sentry self-hosted: FSL-1.1-Apache-2.0, which is not OSI open source (repo).
Commercial SaaS¶
SaaS vendors bill on volume (hosts, GB ingested, spans), while Micromegas runs on your own object storage, so the comparison is a cost model rather than a feature list. The vs. SaaS Vendors pages (methodology, Datadog, Dynatrace, Elastic, Grafana Cloud, New Relic, Splunk) work through the numbers.
Summary: which one fits¶
These projects overlap more than they compete, and many teams run two of them.
Choose Micromegas when efficiency matters, meaning instrumentation overhead and cost:
- you instrument native code and want detailed spans (Rust crates, Unreal plugin), logs and metrics (also from any app through the C ABI, as the Blender add-on does) left on in production;
- you want very high-frequency, high-resolution telemetry, or full-resolution traces without sampling: Rust CPU traces record every span unsampled in production, enabled with
MICROMEGAS_ENABLE_CPU_TRACING=true(lib.rs), andtelemetry.spans.alldoes the same in Unreal, whose default keeps blocks around frame spikes (console variables,SamplingController.h); - your cost depends on the data more than the code (assets, URLs, scripts, queries going through an interpreter or resolver) and you need to know which input was slow, not just which function;
- your telemetry comes from many processes that aren't classic services: desktop or mobile clients, edge devices, batch jobs, CI runners, game clients and servers;
- you need high event volume and long retention at a predictable cost, stored as Parquet on your own object storage, with spans processed only when queried (on-demand processing); one production deployment on AWS runs at about $1,100/month for 449 billion events over 90 days (cost breakdown);
- you want one SQL surface across logs, metrics and traces, including in notebooks, instead of one query language per signal;
- you need per-row access control so each team sees the telemetry meant for it, with privacy guarantees.
Look elsewhere if:
- you want the established, widely adopted default, with the largest community and integration ecosystem (Grafana LGTM, Prometheus, SigNoz);
- you don't want to operate the stack (a SaaS vendor, or a peer's hosted offering);
- you can't run PostgreSQL, need a built-in alert engine, or have PromQL dashboards and alert rules you want to keep.
FAQ¶
What is an open-source, self-hosted alternative to Datadog that I can query with SQL?¶
It depends on the workload. For OTel-instrumented services with a ready-made APM UI, look at SigNoz or ClickStack. For native code and client fleets where efficiency matters, Micromegas is Apache-2.0, self-hosted and queried with SQL over FlightSQL (query guide).
How do I reduce observability costs at high event volume?¶
Micromegas stores raw data in your own object storage and processes spans only when they are queried; one production deployment runs at about $1,100/month for 449 billion events over 90 days (cost breakdown). VictoriaMetrics is a strong choice for metrics, with low resource use and no external dependencies (VictoriaMetrics).
How do I record full-resolution traces in production without sampling?¶
In Rust, set MICROMEGAS_ENABLE_CPU_TRACING=true and every CPU span is recorded unsampled (lib.rs). In Unreal Engine, set telemetry.spans.all 1 (console variables); by default Unreal keeps blocks around frame spikes.
How do I collect telemetry from Unreal Engine games in production?¶
Micromegas has an Unreal Engine plugin that records spans, logs and metrics and ships them to your own object storage. Unreal Insights remains the right tool for a deep look at a single session.
How do I collect telemetry from desktop apps or game clients across many users?¶
Micromegas keeps one row per emission with process, computer and user columns, so SQL can group by any of them across the fleet (measures). Use the Rust or Unreal SDKs, the C ABI for logs and metrics, or OTLP/HTTP.
How do I find which asset, script or query made my code slow, not just which function?¶
Name spans by the data being processed (an FName in Unreal, a statically allocated string in Rust) and attach context as a property set, then query the spans in SQL. A sampling profiler shows the hot function; data-named spans show the input behind it (code vs. data).
What is a Prometheus alternative for sub-second, high-frequency metrics?¶
Prometheus's scrape_interval defaults to 1m (config). Micromegas stores every metric emission as its own row (measures), and GreptimeDB is a good fit for high-cardinality metrics and Prometheus long-term storage.
How do I give each team access to the right telemetry, with privacy guarantees, in a self-hosted observability store?¶
Micromegas stamps every row with an audience taken from the ingestion credential, and data stamped for an audience is visible only to principals granted that audience, so each team sees the data meant for it. All of this is in the Apache-2.0 build (authorization). Grafana Enterprise Logs offers label-based access control, but it is a paid tier (GEL).
How do I drill down from a fleet-wide dashboard to one process without scanning terabytes of data?¶
Define a reduced view with CREATE MATERIALIZED VIEW (for example, per-minute counts by process): Micromegas refreshes it every second and merges it into minute, hour and day partitions, so the fleet-wide query stays small, and the drill-down reads full-resolution rows only for the selected processes and time range (derived views). Old data is not downsampled: raw events keep full resolution for the whole retention period. Prometheus recording rules precompute metrics the same way, for PromQL (precomputation).
How do I trace Rust applications in production with low overhead?¶
Use the micromegas-tracing macros such as span_scope! and #[span_fn], which record spans in-process on the calling thread; existing tracing events are captured as logs (tracing_interop.rs).