Checks from independent regions. A failure confirmed before it is believed. Then, inside the limits you set, the traffic moves off the fault — and the fix is verified before the incident is closed.
No card, no billing yet. One agent, one minute, or no agent at all.
Step through what SystemsWarden actually did — the evidence it gathered, the target it chose and why, and how it proved the fix held.
Frankfurt saw the node fail. Amsterdam and Warsaw were asked before anyone was woken.
Eleven candidates were scored. Three were excluded before scoring even began, because policy said so.
DNS rotated and proxy traffic rebalanced. The log entry was written before the change was applied, not after.
The loop is not closed when the traffic moves. It is closed when the move is proven — or reversed.
Every uptime tool can tell you a server stopped answering. These are the things that happen next.
Degraded infrastructure stops receiving traffic immediately, without being torn out of the fleet — so it can come back when it recovers.
Candidates ranked on spare capacity, live probe results, cooldown state, region and the policy rules you wrote — not on alphabetical order.
Affected domains are distributed across several healthy targets, so recovering from one outage does not manufacture the next one.
Bandwidth and CPU trends are projected forward, so you get told to provision days before the ceiling turns into downtime.
After every action the affected services are re-checked from all regions. If the fix did not hold, the change is rolled back and a human is called.
Nothing changes state without an entry in the action log and a notification. Every domain can answer "why am I on this node, and since when".
Automatic action is only safe if the reasoning is visible, the limits are yours, and the result is checked. Here is the whole loop.
One region seeing a timeout proves something about that region, not about your server.
Targets your policy forbids are never considered, however healthy they look.
Rotate DNS off the failure, rebalance proxy traffic, drop endpoints that stopped serving.
The incident is not resolved because traffic moved. It is resolved when the move is proven.
Detection, decision transparency and the action log are on every plan. Letting SystemsWarden execute the action is the Autopilot plan — and it starts in shadow mode, publishing what it would have done, for as long as you want to read it before you hand over the keys.
The ordinary work, done properly, so you do not need a second tool beside this one.
Status, response time and content checks from multiple regions, as often as every 30 seconds.
Raw port reachability and real SMTP conversations, so a silently dead mail server cannot hide.
A job that stops running is an outage nobody sees. Give it a heartbeat URL and its silence becomes an alert.
CPU, memory, disk, network, I/O wait, steal time, temperature and RAID health — from one small binary that updates itself.
Certificates and domain renewals tracked well before they lapse, because both fail at the worst possible hour.
IP reputation monitoring, and branded status pages you can point your own customers at.
Indicative pricing — nothing is charged during early access. Watching is cheap. Deciding is worth more. Acting on production infrastructure — and standing behind the result — is the part that replaces a person on call.
A personal site or a first server.
Sites whose downtime costs you money.
You run servers, not just websites.
The loop closes. It drains, rotates, verifies and rolls back on its own.
Multi-tenant accounts for hosts and managed service providers, private probe locations inside your own network, import of your existing provider inventory, SLA and dedicated support.
Custom, from $999 / month
Prices in USD, billed monthly, cancel whenever. Every plan — including Free — gets multi-region checking, the decision reasoning and the full action log. What you pay for is how much of the work is taken off your hands.
Not unless you turn that on. Execution is the Autopilot plan and it begins in shadow mode, where it publishes the action it would have taken and does nothing at all. You read those proposals for as long as you like. Even armed, it stays inside your limits — nodes without headroom, nodes in cooldown, and nodes your policy excludes are never chosen. See Security for how permissions and credentials are handled.
After every action the affected services are re-checked from all regions. If they are not serving, or the error rate has not returned to baseline, the change is rolled back to the previous map and a human is escalated to. An action that cannot be verified is not treated as a success.
Every check runs from several independent regions and a service is only marked down when they agree. On top of that, three consecutive confirmed failures are required before anything is acted on, and three clean runs before a recovery is accepted. Across 262 consecutive multi-region probes in our own production fleet, that produced zero false alarms.
For sites, APIs, ports, mail and heartbeats: nothing. To watch a server's CPU, memory, disk and network you drop in one small Go binary. It measured 0.034% CPU and 11.6 MB of memory on a production server, and it updates itself — you never touch it again. You can revoke it from the dashboard at any time.
Yes, and this is deliberate: nothing changes state without writing to the action log and sending a notification. You get a global history and a per-domain history, so "why is this domain on that node, and since when" always has an answer with a timestamp on it.
Nothing moves. Remediation is opt-in and fail-closed: with no confirmed multi-region verdict, no action is taken, and your existing DNS and proxy configuration keeps serving exactly as it was. Detection runs from edge locations independent of the control plane, so a control plane outage cannot invent a false failure either.
Not yet, and we would rather say so than take your card. SystemsWarden is in early access: accounts are free, billing is not switched on, and the prices on this page show where each plan is heading rather than what you will be charged today. Early-access users keep the limits we agree when billing does open.
Email and Telegram today, with webhooks, Slack, Discord and SMS on the roadmap. Delays and thresholds are yours to set, so a two-minute blip on a staging box does not have to reach your phone.
Fleet includes white-label status pages, so your customers see your brand rather than ours, along with the uptime history that backs up whatever you promised them.
Early-access accounts are free, keep the plan limits we agree, and shape what gets built next. Tell us what you run and we will get you set up.