Trabant Systems · IFA Trabant
Five production watchdogs, a stable reference node and the transition from a working installation to a reproducible system.
This note continues the production testing documented in
IFA Trabant: A Second Extended Run on the 2 GB Reference Node
.
The previous test ended with a specific question:
Can several autonomous watchdogs share the same small Linux node while preserving operational isolation?
At that stage, the reference HP Stream had already demonstrated that a persistent monitoring and failover function could remain active for several days without imposing significant computational load on the host.
The next proposed experiment was therefore replication rather than expansion: additional domains, independent watchdogs, independent state and no centralised decision layer.
That experiment is no longer prospective.
The reference node is now running five independent production watchdogs.
The next problem is different.
The watchdog did not become larger.
It became five watchdogs.
The next task is to make the system reproducible.
1. From replication test to production state
The multi-domain architecture follows the same principle used in the original implementation:
Local failure, local response.
Each protected domain retains its own monitoring process and its own operational state.
A failure affecting one domain should not automatically modify the state of another. The watchdogs share a physical machine and a common architectural model, but they do not share a decision state.
This preserves the distinction established during the earlier tests:
Physical Linux node
|
+-- Watchdog A
| independent state
|
+-- Watchdog B
| independent state
|
+-- Watchdog C
| independent state
|
+-- Watchdog D
| independent state
|
+-- Watchdog E
independent state
The relevant result is not that five shell scripts can run on a computer.
That would be trivial.
The relevant observation is that replication has not required the architecture to become centralised.
The node remains simple.
The watchdogs remain autonomous.
The physical host is shared.
The uncertainty is not.
2. Simplifying the network layer
A separate problem emerged during the development period.
The reference antiX installation had accumulated several mechanisms capable of participating in wireless network configuration at different times, including ConnMan, Ceni-generated interface configuration, dhcpcd and wpa_supplicant.
This was undesirable for a machine whose purpose is continuity.
A failover node should not need several competing mechanisms merely to remain connected to the network.
The network layer was therefore simplified around NetworkManager.
NetworkManager now manages the Wi-Fi connection profile, DHCP, routing and DNS. The previous WLAN configuration was removed from /etc/network/interfaces; ConnMan was disabled; the independent dhcpcd service was also removed from the active control path. wpa_supplicant may remain present as the wireless backend used by NetworkManager.
The resulting separation is deliberately uncomplicated:
NetworkManager
|
+-- connection profiles
+-- DHCP
+-- routes
+-- DNS
|
+-- wpa_supplicant
|
Wi-Fi
IFA does not need to become a network manager.
NetworkManager manages connectivity.
IFA evaluates continuity.
That distinction reduces the number of components capable of changing the same state.
3. The physical reference node
The production node remains the same small HP Stream used in the previous extended tests.
It is approximately a 2014-generation machine with 2 GB of installed RAM and roughly 32 GB of internal storage.
The computer was not purchased for IFA.
It was reused.
The previous Windows installation was removed, the machine was reformatted and antiX Linux was installed as the operating system.
The original battery had degraded and was replaced before the Stream was placed into continuous service.
This is relevant to the broader idea of hardware reuse.
An old notebook that has spent several years unused may still be entirely adequate for a narrowly defined infrastructure function, but age alone should not determine suitability.
The machine should first be inspected.
Battery condition deserves particular attention. The battery is not required for IFA to function, but a degraded notebook battery is an unnecessary risk in a device intended to remain powered for long periods.
A healthy replacement battery also has a secondary advantage.
During a short electrical interruption, the notebook can continue operating from its internal battery without requiring an additional device. It therefore provides a limited built-in power reserve.
That does not convert the Stream into redundant infrastructure.
It simply uses a capability the notebook already possesses.
4. Passive thermal management
Continuous operation also raised a simple physical question.
The Stream was designed as a small notebook, not as rack-mounted infrastructure.
No active external cooling has been added.
Instead, the machine is normally left physically open, while the display can remain switched off, and the chassis is elevated slightly above the surface on which it rests.
Three insulated spacers maintain an air gap beneath most of the notebook.
In the production prototype those spacers happen to be three large batteries fitted with protective covers.
Their chemistry and electrical function are irrelevant.
They are being used as three objects of approximately the correct height.
The important element is the space underneath the chassis.
When that space is removed and the notebook rests directly on a surface, the machine becomes warmer, as would be expected with a small passively cooled computer.
With the air gap maintained, the thermal behaviour of the reference node has remained modest during ordinary unattended operation.
In one representative near-simultaneous measurement, after approximately 30 hours of continuous operation, Conky reported an SoC temperature of 35°C, while a separate room thermometer indicated an ambient temperature of 24.3°C. At the same time, the node was using approximately 279 MB of RAM with no swap activity.
This places the reported SoC temperature roughly 11°C above the measured room temperature under that particular operating condition.
These figures are observations from this specific production node rather than thermal specifications for HP Stream hardware generally.
No external USB fan has so far been necessary.
For this particular workload, passive airflow has therefore been preferable: it introduces no motor, no additional cable, no additional noise and no additional mechanical failure point.
The objective is not to obtain the lowest possible temperature.
It is to maintain an adequate and stable thermal regime with as little additional infrastructure as possible.
5. Dedicated operation remains a small workload
The distinction observed during the previous extended run remains important.
IFA itself is not computationally demanding.
During unattended operation, the Stream typically remains around the low hundreds of megabytes of memory usage, with little or no swap activity.
Interactive maintenance is different.
Opening Firefox, web applications, terminals and other contemporary desktop tools may increase memory consumption substantially.
That is not a contradiction.
They are different workloads.
The production function is approximately:
observe
wait
observe
compare
decide
log
wait
Interactive administration temporarily asks the same machine to behave as a modern desktop.
The Stream can do both.
It is simply much better suited to the first.
6. A production freeze
The architecture has now reached a useful point.
It works.
That changes the development problem.
Continuing to modify the production machine without preserving a known working state would gradually create another form of fragility.
The next phase should therefore begin with a production freeze.
This does not mean that IFA is complete.
It means that the existing working configuration should become a reference version before further abstraction begins.
The distinction is important:
working installation
!=
reproducible system
At present, the Stream contains both the architecture and the history of its development.
The next step is to separate them.
7. Removing what the node does not need
The current antiX installation still contains applications that were useful while the machine was being developed as a general Linux computer but do not contribute to its permanent role.
LibreOffice is one obvious example.
A dedicated ChatGPT client is another.
Neither is required for monitoring, failover or ordinary system administration.
They can therefore be removed before the reference system is packaged.
This should not become an exercise in producing the smallest possible Linux distribution.
The objective is more conservative.
Software with no meaningful role should disappear.
Useful general tools should remain.
The intended base therefore continues to include NetworkManager and nmtui, wpa_supplicant, a lightweight browser, terminal, text editor, Conky, ordinary network utilities, the IFA components and the antiX tools required to preserve Live and installation functionality.
The result should still look recognisably like a Linux system.
Just one with a defined job.
8. Recovery before distribution
The first image produced from the reference node should not be the public version.
It should be the recovery version.
The purpose of Trabant IFA Recovery is practical.
If the existing Stream fails tomorrow, the architecture should not have to be reconstructed manually from notes, memory and individual scripts.
A replacement compatible computer should be capable of booting from USB, running the environment as a Live system and, where appropriate, installing it onto local storage.
The recovery target is therefore approximately:
failed node
|
replacement PC
|
boot Trabant IFA Recovery
|
verify hardware and network
|
install if required
|
restore production operation
The intended recovery time is measured in tens of minutes rather than in a new development session.
This private image may contain the actual production configuration required by the system.
Secrets, however, require separate treatment.
A recovery mechanism should preserve operational capability without turning a misplaced USB device into an uncontrolled copy of production credentials.
9. From Recovery to Open
Only after the Recovery image has been created and tested on different hardware should a generic public edition be derived.
Trabant IFA Open should use the same architecture and the same code.
It should not contain the same configuration.
The public system does not need a graphical setup assistant.
It is not intended to hide Linux from its user.
A technically competent user should instead be able to open documented configuration files and replace clearly identified values.
For example:
DOMAIN="YOUR_DOMAIN_HERE"
PRIMARY_CONNECTION="YOUR_PRIMARY_CONNECTION"
FAILOVER_CONNECTION="YOUR_FAILOVER_CONNECTION"
Cloudflare-related parameters can be kept in a separate configuration file.
Credentials should be isolated from both the scripts and the ordinary configuration.
The public image can therefore distribute templates such as:
ifa.conf.example
cloudflare.conf.example
secrets.conf.example
with accompanying documentation explaining what each variable means and where it must be changed.
No setup wizard is required.
The intended user is not the general public.
The intended user is someone comfortable enough with Linux to read a text file, understand a domain name, configure a network profile and insert an API credential in the documented location.
For that user, transparency is more useful than another interface.
10. Configuration should not become code
Before a public image can exist, the current scripts must therefore be audited.
Production-specific values should not remain embedded in executable logic.
The intended separation is conceptually simple:
/usr/local/lib/trabant-ifa/
program logic
/etc/trabant-ifa/
configuration
/etc/trabant-ifa/secrets.conf
credentials
/usr/share/doc/trabant-ifa/
documentation
The code describes how IFA behaves.
The configuration describes where it operates.
The secrets authorise actions.
Those are different concerns.
They should remain different files.
The public version should also detect obvious unresolved placeholders rather than silently start with values such as:
YOUR_DOMAIN_HERE
YOUR_CLOUDFLARE_API_TOKEN
A configuration error should fail visibly.
Silent ambiguity is less useful.
11. The Stream should not become part of the specification
There is one further abstraction to make.
The current machine has a particular wireless interface name.
Another computer may not.
Depending on hardware and Linux naming conventions, an adapter might appear as:
wlan0
wlan1
wlp2s0
A public IFA implementation should therefore avoid assuming that a particular physical interface exists.
Where possible, connectivity should be described through NetworkManager profiles and discovered state rather than hard-coded adapter names.
The system should be capable of identifying available connections, active routes and the selected primary or failover profile without requiring the replacement machine to reproduce the Stream’s hardware naming.
The HP Stream is the current reference host.
It should not become an undocumented software dependency.
12. The next phase
The development sequence is now relatively clear.
First, leave the present NetworkManager configuration running long enough to confirm that the simplification has not introduced a new connectivity problem.
Then freeze the current production state.
Audit the scripts.
Remove embedded production assumptions.
Separate code, configuration and secrets.
Remove unnecessary desktop software.
Create the private Recovery image.
Boot that image on another compatible computer.
Test hardware detection.
Test NetworkManager.
Test the watchdogs.
Test Live operation.
Test installation.
Only then derive the generic Open edition.
The sequence can be reduced to:
production
|
v
freeze
|
v
audit
|
v
clean
|
v
Recovery image
|
v
different hardware
|
v
generic configuration
|
v
Open release
Conclusion
The previous IFA Trabant production note ended by asking whether several autonomous watchdogs could share the same small Linux node while preserving operational isolation.
They now do.
Five production watchdogs share the reference HP Stream without requiring a central decision state.
The interesting development problem has therefore moved again.
It is no longer primarily endurance.
It is no longer replication.
It is preservation.
A working installation is useful.
A working installation that can be reproduced after the original machine disappears is more useful.
And a reproducible system whose production-specific assumptions can subsequently be removed becomes something else again:
a system that can be shared.
The next experiment is therefore not another watchdog.
It is whether Trabant IFA can survive the loss of the Trabant.