Pre-1.0 hoardDB is pre-1.0. Expect breaking changes.
Is this node full?
A node runs out of one of four things: file descriptors, memory mappings, memory, or disk. hoardDB measures all four against the limits your kernel and your filesystem actually enforce, and tells you which one is close. Read which resource is red before you act: the right action differs by resource, and for disk the usual reflex, adding a node, is the wrong one.
The one-sentence version: look at which resource is red. For fds or
mappings, raise the kernel limit. For memory under load, add a node. For
disk, free space on this node: a red node votes against new members, but in a
cluster of three or more the others can still admit one, and a new node barely
helps a full disk either way.
Where to look
You do not need to set anything up to be told. Three channels carry the same readings:
The node’s log. A WARN line is written when a resource crosses its line, and repeated at most once an hour while it stays there. Each line names the action for that resource. This is always on.
hoardDB status. Any authenticated session shows one row per resource:[h]oardDB> status STATUS --- ... Healthy: true Capacity: fds 0.00 threshold 0.80 ok mappings 0.00 threshold 0.80 ok memory 0.00 threshold 0.80 ok disk 0.32 threshold 0.85 okThis is always on too.
Prometheus, at
/metrics. Off by default; see Turning on the gauges.
What the states mean
| state | what it means |
|---|---|
ok | below the threshold |
approaching | disk only, with the high-water mark on: within 0.10 of it. Plan now |
red | at or over the threshold. Act on the row for that resource, below |
unavailable | not fine. The resource could not be measured, so there is no ratio to show. The log says why, and what to check |
stale | not fine. The last reading is more than 45 seconds old: whatever measures the node has stopped. The ratios shown are the last ones known, not current ones |
Neither unavailable nor stale means healthy. A node that cannot be measured
is a node you know nothing about.
A red row in status is followed by one line pointing you to the log, because
the log line carries the action and, for memory, which of two things red means.
What to watch, with Prometheus
hoarddb_node_capacity_ratio{resource} is each resource’s reading divided by
its live limit. hoarddb_node_capacity_threshold_ratio{resource} is where that
resource turns red. The resource label takes exactly the values fds,
mappings, memory and disk. Graph the four ratios with:
max by (resource) (hoarddb_node_capacity_ratio)
Alert on three rules, not one. The first alone is silent when a resource stops being measured, because the ratio then disappears rather than reading zero:
# a resource is over its threshold
hoarddb_node_capacity_ratio > on(resource) hoarddb_node_capacity_threshold_ratio
# a resource can no longer be measured (the threshold is always published; the ratio is not)
hoarddb_node_capacity_threshold_ratio unless on(resource) hoarddb_node_capacity_ratio
# the sampler has stopped
time() - hoarddb_node_capacity_last_sample_timestamp_seconds > 60
The thresholds
The thresholds are policy lines, not measured cliffs. They are where this release chooses to warn you, not points at which hoardDB has been measured to fail. Where your node actually degrades depends on your hardware and workload, and no figure for it is published yet.
| resource | red at | amber |
|---|---|---|
fds | 0.80 | none |
mappings | 0.80 | none |
memory | 0.80 | none |
disk | disk_high_water_pct ÷ 100 (default 0.85), or 0.80 when that is 0 | 0.10 below red, only while disk_high_water_pct is on |
None of these is configurable except through disk_high_water_pct, which
already exists for another reason.
Disk and disk_high_water_pct
server.disk_high_water_pct (HOARDB_DISK_HIGH_WATER, default 85) is the
point at which this node votes against new members. The disk warning fires at
the same point, from the same measurement, so the warning and the vote always
agree.
0turns off only the refusal. The node no longer votes against joins or refuses to propose one, but the disk still warns, at 0.80: turning off the vote does not turn off being told your disk is filling. There is no amber line in this mode, because there is no vote to warn you ahead of.- A value above
100turns off both the refusal and the warning. The threshold is then above anything a disk can reach.
When a resource is red
Read the disk row first. It is the one where the obvious action is wrong.
| resource red | what it usually means | do this | does adding a node help? |
|---|---|---|---|
disk | The data directory’s filesystem is filling up: the store’s own bytes, a write-ahead log held back for a replica that is unreachable, blob payloads, or something that is not hoardDB. | Free space on this node. (1) Compare hoarddb_node_data_bytes with hoarddb_node_disk_used_bytes to see whether hoardDB is the consumer at all. (2) If hoarddb_node_data_bytes{part="wal"} is large, bring the unreachable peer back or remove it from the cluster: the log is kept for it on purpose. (3) If hoarddb_reclaim_pending_records is above 0, dropped data is still being reclaimed, and its space returns only after compaction. (4) Grow the volume, or drop buckets or databases you no longer need. Do not grow the cluster while this node is red. | No. See below. |
disk, approaching | as above, earlier | Plan now. If growth is organic and even across buckets, add capacity while every node is below the mark, so every node votes for the join. If one bucket or the log is growing, treat it as the red row. | Yes, as planning, while every node’s disk is below the mark |
fds | Descriptors are held mostly by the store’s table and value-log files, which grow with stored bytes, plus client and peer connections (server.max_conns and server.max_peer_conns) | Raise the process’s RLIMIT_NOFILE: LimitNOFILE= in a systemd unit, --ulimit nofile= in Docker. It is a kernel limit, and raising it is free. If it is already very high and still close, the node holds a lot of data: read the disk row. | Not as the first action |
mappings | The store maps its table files into memory, so mappings also grow with stored bytes | Raise vm.max_map_count: sysctl -w vm.max_map_count=262144, and persist it in /etc/sysctl.d/. The kernel default of 65530 is low for any database that maps its files. | Not as the first action |
memory | Anonymous memory, resident or swapped: the block and index caches (storage.badger_block_cache_mb and badger_index_cache_mb), per-request working memory, the in-memory bucket registry, and a rebalance in progress, which holds what it moves in memory on both nodes. Memory is red: read the log’s memory-limit line first; it tells you which of two things red means. Against memory.max or the host’s RAM, the node is near where it swaps heavily or is killed, so act now. Against memory.high, the kernel is already throttling this process. The node is slow, not about to be killed. memory.high is a throttle line, not a hard limit | (1) If a rebalance is running, wait for it; the peak passes. (2) If the caches were raised, lower them. (3) If the load is traffic-driven, add a node, and expect a brief rise on this node while buckets move away. (4) Against memory.max or host RAM: if the limit is below the host’s memory, raise it. Against memory.high: raise it if it was set lower than you meant (MemoryHigh= in systemd), and add a node if the load is real. | Yes, when traffic or the number of buckets drives it. It does not shrink the fixed caches, and the rebalance itself costs memory first |
Why adding a node does not fix a full disk
Disk is red: free space on this node. Adding a node is not the fix. At the high-water mark this node votes against new members. In a one- or two-node cluster that blocks the join. In a cluster of three or more, the other nodes can still admit one, and this node then takes part in the rebalance, so do not grow the cluster while a node’s disk is red. Either way, a new node frees little space on this one, and slowly: buckets move by hash and not by size, freed space comes back only after cleanup and compaction, and space held by values of 1 MiB or more is not returned at all in this release. Add nodes while every node’s disk is green, as planning.
The log line for a red disk says the same:
disk on this node is at 0.91 of its filesystem (high-water mark 0.85): free space on this node or grow the volume. At this mark the node votes against joins, but in a cluster of three or more a majority can still admit one, and this node then takes part in the rebalance. Do not grow the cluster while this node is red
and, with disk_high_water_pct on, the early warning:
disk on this node is at 0.78, approaching the high-water mark 0.85: free space or grow the volume now. At the mark this node votes against joins
What the memory ratio counts
The memory ratio is anonymous memory, resident plus swapped (RssAnon + VmSwap
from /proc/<pid>/status), over the memory limit. Page cache does not count:
the kernel sheds it under pressure.
Swapped memory counts as used, so a node that is swapping reads high and can
read above 1.0. Against memory.max or host RAM, above 1.0 means the node is
over its hard limit and swapping. Against memory.high, it means the node is
past the line where the kernel throttles it.
The limit is the smallest memory.max or memory.high on the process’s cgroup
and every parent cgroup it can see, or the host’s MemTotal when none is set.
The log says which one it is using, when the node starts and again whenever it
changes:
memory limit 4.0 GiB from memory.max on cgroup /system.slice
memory limit 2.0 GiB from memory.high on cgroup /system.slice/hoarddb.service: this is a throttle line, not a hard limit. Past it the kernel slows this process and reclaims its memory; it does not kill it there. The hard limit is 4.0 GiB from memory.max on cgroup /system.slice
The two red lines read differently, on purpose:
memory on this node is at 0.86 of its memory limit 4.0 GiB (memory.max on cgroup /system.slice): anonymous memory, resident plus swapped, is near the point where this process swaps heavily or is killed. If a rebalance is running, wait for it. If the block or index caches were raised, lower them. If the load is traffic-driven, add a node. If the limit is below the host's memory, raise it.
memory on this node is at 0.86 of memory.high 2.0 GiB (cgroup /system.slice/hoarddb.service), the line where the kernel starts throttling this process and reclaiming its memory. The node is being slowed, not stopped: memory.high is not a hard limit, and the process is not killed there. If memory.high was set lower than you meant (MemoryHigh= in systemd), raise it. If a rebalance is running, wait for it. If the block or index caches were raised, lower them. If the load is traffic-driven, add a node.
Inside a container, set the memory limit on the container itself. A process
in its own cgroup namespace, which is Docker’s default, cannot see a limit set
above the container, for example only on a pod. The log then says the limit in
use is the host’s MemTotal and that a limit above the container is not seen.
On a host that manages memory with cgroup v1, hoardDB does not read the v1
limit. It uses the host’s MemTotal and writes a warning at start and every
hour saying so; if the process runs under a v1 memory limit, the memory ratio
reads lower than it should.
Turning on the gauges
The gauges are exported only when /metrics is on. It is a deploy-time
setting:
HOARDB_METRICS_ENABLED=true HOARDB_METRICS_LISTEN=10.0.0.5:9090 hoardDB-server
The log warnings and status work without it. When it is off, the node says so
once at start:
capacity gauges are not exported (HOARDB_METRICS_ENABLED is false); capacity warnings still go to this log and to hoardDB status
/metrics is unauthenticated. It listens on 0.0.0.0:9090 by default. Bind
it to a private interface, as above, or firewall the port.
The series, all node-wide (none is per bucket, per database or per peer):
| series | what it is |
|---|---|
hoarddb_node_capacity_ratio{resource} | reading ÷ limit. Absent while the resource cannot be measured |
hoarddb_node_capacity_threshold_ratio{resource} | where that resource turns red. Always present |
hoarddb_node_capacity_sample_errors_total{resource} | samples in which the resource could not be read |
hoarddb_node_capacity_last_sample_timestamp_seconds | when the last sample completed |
hoarddb_node_buckets | buckets this node serves |
hoarddb_node_memory_mappings, hoarddb_node_memory_mappings_limit | mappings, and vm.max_map_count |
hoarddb_node_memory_anon_bytes, hoarddb_node_memory_limit_bytes | the memory ratio’s two sides |
hoarddb_node_disk_used_bytes, hoarddb_node_disk_size_bytes | the disk ratio’s two sides, from one measurement |
hoarddb_node_data_bytes{part} | what hoardDB itself holds: store_lsm, store_vlog (up to a minute old) and wal. wal is absent on a single node, which keeps no log |
process_open_fds, process_max_fds | the descriptor ratio’s two sides |
Reading your node’s limits
Every ratio divides by a limit you can read yourself. Replace <pid> with the
server’s process id (pgrep hoardDB-server) and <data_dir> with its data
directory.
# descriptors: the "Max open files" soft limit is the fds limit
grep 'Max open files' /proc/<pid>/limits
ls /proc/<pid>/fd | wc -l
# mappings
sysctl vm.max_map_count
wc -l < /proc/<pid>/maps
# disk
df -h <data_dir>
# memory: the limits on the process's cgroup and every parent it can see
cg=$(sed -n 's/^0:://p' /proc/<pid>/cgroup)
while [ -n "$cg" ]; do echo "$cg memory.max=$(cat /sys/fs/cgroup$cg/memory.max 2>/dev/null) memory.high=$(cat /sys/fs/cgroup$cg/memory.high 2>/dev/null)"; cg=${cg%/*}; done
grep -E 'RssAnon|VmSwap' /proc/<pid>/status
max means no limit at that level. The walk assumes cgroup v2 mounted at
/sys/fs/cgroup, which is the common layout; hoardDB itself finds the mount
through /proc/<pid>/mountinfo.
Not in the capacity set
- CPU is a watch signal, not a ratio:
rate(process_cpu_seconds_total[5m])against the node’s core count. Sustained, request-driven CPU is the classic case where adding a node helps. CPU during a rebalance, a reclaim or a compaction is expected, and passes. process_network_receive_bytes_totalandprocess_network_transmit_bytes_totaldo not measure hoardDB. They read the network namespace’s counters, so on a host with shared networking they report the whole machine’s traffic. Usehoarddb_request_duration_secondsto see whether the network is slowing requests down.
Source: docs/user/capacity.md in the repository.