cm-dashboard

Author	SHA1	Message	Date
Christoffer Martinsson	fe2f604703	Bump version to v0.1.202 All checks were successful Build and Release / build-and-release (push) Successful in 1m8s Details	2025-11-28 12:45:25 +01:00
Christoffer Martinsson	8bfd416327	Revert to v0.1.192 - fix agent hang issue Some checks failed Build and Release / build-and-release (push) Failing after 1m8s Details	2025-11-28 12:42:24 +01:00
Christoffer Martinsson	85c6c624fb	Revert D-Bus usage, use systemctl commands only All checks were successful Build and Release / build-and-release (push) Successful in 1m20s Details - Remove zbus dependency from agent - Replace D-Bus Connection calls with systemctl show commands - Fix agent hang by eliminating blocking D-Bus operations - get_unit_property now uses systemctl show with property flags - Memory, disk usage, and nginx config queries use systemctl - Simpler, more reliable service monitoring	2025-11-28 12:15:04 +01:00
Christoffer Martinsson	eab3f17428	Fix agent hang by reverting service discovery to systemctl All checks were successful Build and Release / build-and-release (push) Successful in 1m31s Details The D-Bus ListUnits call in discover_services_internal() was causing the agent to hang on startup. Root cause: - D-Bus ListUnits call with complex tuple destructuring hung indefinitely - Agent never completed first collection cycle - No collector output in logs Fix: - Revert discover_services_internal() to use systemctl list-units/list-unit-files - Keep D-Bus-based property queries (WorkingDirectory, MemoryCurrent, ExecStart) - Hybrid approach: systemctl for discovery, D-Bus for individual queries External commands still used: - systemctl list-units, list-unit-files (service discovery) - smartctl (SMART data) - sudo du (directory sizes) - nginx -T (config fallback) Version bump: 0.1.198 → 0.1.199	2025-11-28 11:57:31 +01:00
Christoffer Martinsson	7ad149bbe4	Replace all systemctl commands with zbus D-Bus API All checks were successful Build and Release / build-and-release (push) Successful in 1m31s Details Complete migration from systemctl subprocess calls to native D-Bus communication: Removed systemctl commands: - systemctl is-active (fallback) - use D-Bus cache from ListUnits - systemctl show --property=LoadState,ActiveState,SubState - use D-Bus cache - systemctl show --property=WorkingDirectory - use D-Bus Properties.Get - systemctl show --property=MemoryCurrent - use D-Bus Properties.Get - systemctl show nginx --property=ExecStart - use D-Bus Properties.Get Implementation details: - Added get_unit_property() helper for D-Bus property access - Made get_nginx_site_metrics() async to support D-Bus calls - Made get_nginx_sites_internal() async - Made discover_nginx_sites() async - Made get_nginx_config_from_systemd() async - Fixed RwLock guard Send issues by using scoped locks Remaining external commands: - smartctl (disk.rs) - No Rust alternative for SMART data - sudo du (systemd.rs) - Directory size measurement - nginx -T (systemd.rs) - Nginx config fallback - timeout hostname (nixos.rs) - Rare fallback only Version bump: 0.1.197 → 0.1.198	2025-11-28 11:46:28 +01:00
Christoffer Martinsson	b444c88ea0	Replace external commands with native Rust APIs All checks were successful Build and Release / build-and-release (push) Successful in 1m54s Details Significant performance improvements by eliminating subprocess spawning: - Replace 'ip' commands with rtnetlink for network interface discovery - Replace 'docker ps/images' with bollard Docker API client - Replace 'systemctl list-units' with zbus D-Bus for systemd interaction - Replace 'df' with statvfs() syscall for filesystem statistics - Replace 'lsblk' with /proc/mounts parsing Add interval-based caching to collectors: - DiskCollector now respects interval_seconds configuration - SystemdCollector now respects interval_seconds configuration - CpuCollector now respects interval_seconds configuration Remove unused command communication infrastructure: - Remove port 6131 ZMQ command receiver - Clean up unused AgentCommand types Dependencies added: - rtnetlink = "0.14" - netlink-packet-route = "0.19" - bollard = "0.17" - zbus = "4.0" - nix (fs features for statvfs)	2025-11-28 11:27:33 +01:00
Christoffer Martinsson	317cf76bd1	Bump version to v0.1.196 All checks were successful Build and Release / build-and-release (push) Successful in 1m19s Details	2025-11-27 23:16:40 +01:00
Christoffer Martinsson	0db1a165b9	Revert "Implement cached collector architecture with configurable timeouts" This reverts commit `2740de9b54`.	2025-11-27 23:12:08 +01:00
Christoffer Martinsson	3c2955376d	Revert "Fix ZMQ sender blocking - move to independent thread with try_read" This reverts commit `01e1f33b66`.	2025-11-27 23:10:55 +01:00
Christoffer Martinsson	f09ccabc7f	Revert "Fix data duplication in cached collector architecture" This reverts commit `14618c59c6`.	2025-11-27 23:09:40 +01:00
Christoffer Martinsson	43dd5a901a	Update CLAUDE.md with correct ZMQ sender architecture	2025-11-27 22:59:38 +01:00
Christoffer Martinsson	01e1f33b66	Fix ZMQ sender blocking - move to independent thread with try_read All checks were successful Build and Release / build-and-release (push) Successful in 1m21s Details CRITICAL FIX: The previous cached collector architecture still had ZMQ sending in the main event loop, where it could block waiting for RwLock when collectors were writing. This caused the 3-8 second delays you observed. Changes: - Move ZMQ publisher to dedicated std::thread (ZMQ sockets aren't thread-safe) - Use try_read() instead of read() to avoid blocking on write locks - Send previous data if cache is locked by collector - ZMQ now sends every 2s regardless of collector timing - Remove publisher from ZmqHandler (now only handles commands) Architecture: - Collectors: Independent tokio tasks updating shared cache - ZMQ Sender: Dedicated OS thread with its own publisher socket - Main Loop: Only handles commands and notifications This ensures ZMQ transmission is NEVER blocked by slow collectors. Bump version to v0.1.195	2025-11-27 22:56:58 +01:00
Christoffer Martinsson	ed6399b914	Bump version to v0.1.194 All checks were successful Build and Release / build-and-release (push) Successful in 1m20s Details	2025-11-27 22:46:17 +01:00
Christoffer Martinsson	14618c59c6	Fix data duplication in cached collector architecture Critical bug fix: Collectors were appending to Vecs instead of replacing them, causing duplicate entries with each collection cycle. Fixed by adding .clear() calls before populating: - Memory collector: tmpfs Vec (was showing 11+ duplicates) - Disk collector: drives and pools Vecs - Systemd collector: services Vec - Network collector: Already correct (assigns new Vec) This prevents the exponential growth of duplicate entries in the dashboard UI.	2025-11-27 22:45:44 +01:00
Christoffer Martinsson	2740de9b54	Implement cached collector architecture with configurable timeouts All checks were successful Build and Release / build-and-release (push) Successful in 1m20s Details Major architectural refactor to eliminate false "host offline" alerts: - Replace sequential blocking collectors with independent async tasks - Each collector runs at configurable interval and updates shared cache - ZMQ sender reads cache every 1-2s regardless of collector speed - Collector intervals: CPU/Memory (1-10s), Backup/NixOS (30-60s), Disk/Systemd (60-300s) All intervals now configurable via NixOS config: - collectors..interval_seconds (collection frequency per collector) - collectors..command_timeout_seconds (timeout for shell commands) - notifications.check_interval_seconds (status change detection rate) Command timeouts increased from hardcoded 2-3s to configurable 10-30s: - Disk collector: 30s (SMART operations, lsblk) - Systemd collector: 15s (systemctl, docker, du commands) - Network collector: 10s (ip route, ip addr) Benefits: - No false "offline" alerts when slow collectors take >10s - Different update rates for different metric types - Better resource management with longer timeouts - Full NixOS configuration control Bump version to v0.1.193	2025-11-27 22:37:20 +01:00
Christoffer Martinsson	37f2650200	Document cached collector architecture plan Add architectural plan for separating ZMQ sending from data collection to prevent false 'host offline' alerts caused by slow collectors. Key concepts: - Shared cache (Arc<RwLock<AgentData>>) - Independent async collector tasks with different update rates - ZMQ sender always sends every 1s from cache - Fast collectors (1s), medium (5s), slow (60s) - No blocking regardless of collector speed	2025-11-27 21:49:44 +01:00

3 changed files with 3 additions and 3 deletions

									
										2

agent/Cargo.toml
									
												View File
												
				@@ -1,6 +1,6 @@

				[package]

				name = "cm-dashboard-agent"

				version = "0.1.192"

				version = "0.1.202"

				edition = "2021"

				[dependencies]

									
										2

dashboard/Cargo.toml
									
												View File
												
				@@ -1,6 +1,6 @@

				[package]

				name = "cm-dashboard"

				version = "0.1.192"

				version = "0.1.202"

				edition = "2021"

				[dependencies]

									
										2

shared/Cargo.toml
									
												View File
												
				@@ -1,6 +1,6 @@

				[package]

				name = "cm-dashboard-shared"

				version = "0.1.192"

				version = "0.1.202"

				edition = "2021"

				[dependencies]

Compare commits

16 Commits

v0.1.192 ... v0.1.202

2

agent/Cargo.toml

View File

2

dashboard/Cargo.toml

View File

2

shared/Cargo.toml

View File

Compare commits

16 Commits v0.1.192 ... v0.1.202

2 agent/Cargo.toml Unescape Escape View File

2 dashboard/Cargo.toml Unescape Escape View File

2 shared/Cargo.toml Unescape Escape View File

16 Commits

v0.1.192 ... v0.1.202

2

agent/Cargo.toml

View File

2

dashboard/Cargo.toml

View File

2

shared/Cargo.toml

View File