Early access. SystemsWarden is in active development — accounts are free while we build, and billing is not open yet. Prices below show where each plan is heading.
Closed-loop infrastructure operations

Most monitors tell you what broke.
This one tells you what it did about it.

Checks from independent regions. A failure confirmed before it is believed. Then, inside the limits you set, the traffic moves off the fault — and the fix is verified before the incident is closed.

No card, no billing yet. One agent, one minute, or no agent at all.

0.034%
Agent CPU on a monitored server, measured under load
11.6 MB
Resident memory for the whole agent
0
False alarms across 262 consecutive multi-region probes
72s
Detected to verified, in the incident replayed below
Incident replay

One node degraded at 03:14. Nobody was awake for it.

Step through what SystemsWarden actually did — the evidence it gathered, the target it chose and why, and how it proved the fix held.

incident #4471 · proxy node ams-04 · 03:14:02 UTC resolved automatically

A timeout in one region proves nothing. Three regions voting proves a lot.

Frankfurt saw the node fail. Amsterdam and Warsaw were asked before anyone was woken.

  • First failure03:14:02 — Frankfurt, connect timeout
  • Region vote2 of 3 down — Frankfurt fail, Warsaw fail, Amsterdam degraded
  • Confirmation rule3 consecutive failed runs before anything moves
  • Node metricsCPU 97% · I/O wait 41% · steal 18% · load 34.2
  • Blast radius8 domains served by this node
Confirmed down at 03:14:31

Not "the next node" — the node your rules allow and your capacity supports.

Eleven candidates were scored. Three were excluded before scoring even began, because policy said so.

  • Excluded by policy3 nodes — jurisdiction rule on 2, drain flag on 1
  • Excluded by cooldown2 nodes — rotated within the last 30 min
  • Excluded by headroom3 nodes — under 25% bandwidth free
  • Chosenfra-02, ams-07, waw-01 — load spread, not stacked
  • Why fra-02 scored top63% headroom · probes 3/3 · cooldown clear · latency +12 ms
Plan ready in 1 second

Drain the failure, move the traffic, write down every step.

DNS rotated and proxy traffic rebalanced. The log entry was written before the change was applied, not after.

  • ActionDrain ams-04, rotate 8 domains across 3 targets
  • Distributionfra-02 4 · ams-07 2 · waw-01 2
  • Applied03:14:33 — DNS TTL 120 s, proxy map updated live
  • RecordedAction log (global + per domain), Telegram, webhook
  • Held backams-04 kept in the fleet, marked draining — not deleted
Traffic moved

An action nobody checked is just a hopeful change.

The loop is not closed when the traffic moves. It is closed when the move is proven — or reversed.

  • Re-checked from3 of 3 regions — Frankfurt, Amsterdam, Warsaw
  • Domains serving8 of 8 — HTTP 200, content check passed
  • Error rateback to baseline within 41 s of the change
  • Target load afterfra-02 headroom 51% — still inside its own limit
  • If it had failedRollback to the previous map, and escalate to a human
Verified · incident closed at 03:15:14
Reconstruction of a real remediation, with node names and domains anonymised.
Beyond monitoring

What SystemsWarden does that a monitor cannot.

Every uptime tool can tell you a server stopped answering. These are the things that happen next.

Drain

Take a failing node out of service

Degraded infrastructure stops receiving traffic immediately, without being torn out of the fleet — so it can come back when it recovers.

Score

Choose the replacement on evidence

Candidates ranked on spare capacity, live probe results, cooldown state, region and the policy rules you wrote — not on alphabetical order.

Rebalance

Spread the load, never stack it

Affected domains are distributed across several healthy targets, so recovering from one outage does not manufacture the next one.

Forecast

See the capacity wall coming

Bandwidth and CPU trends are projected forward, so you get told to provision days before the ceiling turns into downtime.

Verify

Prove the fix, or undo it

After every action the affected services are re-checked from all regions. If the fix did not hold, the change is rolled back and a human is called.

Account

Leave a record nobody can dispute

Nothing changes state without an entry in the action log and a notification. Every domain can answer "why am I on this node, and since when".

Detect → Decide → Act → Verify

An alert is where most tools stop. It is where this one starts.

Automatic action is only safe if the reasoning is visible, the limits are yours, and the result is checked. Here is the whole loop.

01 · Detect

Believe it only when regions agree

One region seeing a timeout proves something about that region, not about your server.

  • Multi-region consensus
  • 3 failures before acting
  • 3 clean runs before restoring
02 · Decide

Score the alternatives against your rules

Targets your policy forbids are never considered, however healthy they look.

  • Skip nodes without headroom
  • Honour per-node cooldowns
  • Spread load, never stack it
  • Respect jurisdiction rules
03 · Act

Move the traffic, then write it down

Rotate DNS off the failure, rebalance proxy traffic, drop endpoints that stopped serving.

  • Logged before it is applied
  • Pushed to your alert channel
  • Shadow mode: proposals only
04 · Verify

Close the loop, or reverse it

The incident is not resolved because traffic moved. It is resolved when the move is proven.

  • Re-check from every region
  • Error rate back to baseline
  • Rollback if it did not hold
  • Escalate when it cannot be fixed

Detection, decision transparency and the action log are on every plan. Letting SystemsWarden execute the action is the Autopilot plan — and it starts in shadow mode, publishing what it would have done, for as long as you want to read it before you hand over the keys.

Coverage

And yes — it monitors everything else too.

The ordinary work, done properly, so you do not need a second tool beside this one.

HTTP / HTTPS

Sites and APIs

Status, response time and content checks from multiple regions, as often as every 30 seconds.

TCP / SMTP

Ports and mail

Raw port reachability and real SMTP conversations, so a silently dead mail server cannot hide.

Heartbeat

Cron and background jobs

A job that stops running is an outage nobody sees. Give it a heartbeat URL and its silence becomes an alert.

Agent

Server resources

CPU, memory, disk, network, I/O wait, steal time, temperature and RAID health — from one small binary that updates itself.

Certificates

SSL and domain expiry

Certificates and domain renewals tracked well before they lapse, because both fail at the worst possible hour.

On the roadmap

Blacklists and public status pages

IP reputation monitoring, and branded status pages you can point your own customers at.

Pricing

Priced by how much of the loop you want closed.

Indicative pricing — nothing is charged during early access. Watching is cheap. Deciding is worth more. Acting on production infrastructure — and standing behind the result — is the part that replaces a person on call.

Free
$0 /mo

A personal site or a first server.

  • 10 monitors, 5-minute checks
  • 7 days of history
  • Email and Telegram alerts
  • SSL and domain expiry
  • 1 status page, 1 user
Join early access
Pro
$19 /mo

Sites whose downtime costs you money.

  • 50 monitors, 60-second checks
  • Multi-region consensus, 2 of N
  • 30 days of history
  • All alert channels
  • 3 status pages, 3 users
Join early access
Fleet
$99 /mo

You run servers, not just websites.

  • 200 monitors, 30-second checks
  • Server agent and live load map
  • Bandwidth caps and capacity forecasting
  • Anomaly detection on your baselines
  • White-label status pages, blacklists
  • 90 days of history, 10 users
Join early access
Autopilot
$299 /mo

The loop closes. It drains, rotates, verifies and rolls back on its own.

  • Everything in Fleet
  • Automatic drain, rotation and rebalancing
  • Post-action verification and rollback
  • Guard rails you define, shadow mode first
  • Predictive alerting before the breach
  • Full action log, per domain and global
Request access
MSP & Enterprise

Many organisations, many fleets, one operator.

Multi-tenant accounts for hosts and managed service providers, private probe locations inside your own network, import of your existing provider inventory, SLA and dedicated support.

Custom, from $999 / month

Talk to us

Prices in USD, billed monthly, cancel whenever. Every plan — including Free — gets multi-region checking, the decision reasoning and the full action log. What you pay for is how much of the work is taken off your hands.

Questions

What operators ask before they trust it.

Will it act on my infrastructure without asking?

Not unless you turn that on. Execution is the Autopilot plan and it begins in shadow mode, where it publishes the action it would have taken and does nothing at all. You read those proposals for as long as you like. Even armed, it stays inside your limits — nodes without headroom, nodes in cooldown, and nodes your policy excludes are never chosen. See Security for how permissions and credentials are handled.

What happens if it makes the wrong call?

After every action the affected services are re-checked from all regions. If they are not serving, or the error rate has not returned to baseline, the change is rolled back to the previous map and a human is escalated to. An action that cannot be verified is not treated as a success.

How do you avoid waking me for nothing?

Every check runs from several independent regions and a service is only marked down when they agree. On top of that, three consecutive confirmed failures are required before anything is acted on, and three clean runs before a recovery is accepted. Across 262 consecutive multi-region probes in our own production fleet, that produced zero false alarms.

What do I have to install?

For sites, APIs, ports, mail and heartbeats: nothing. To watch a server's CPU, memory, disk and network you drop in one small Go binary. It measured 0.034% CPU and 11.6 MB of memory on a production server, and it updates itself — you never touch it again. You can revoke it from the dashboard at any time.

Can I see what it changed, after the fact?

Yes, and this is deliberate: nothing changes state without writing to the action log and sending a notification. You get a global history and a per-domain history, so "why is this domain on that node, and since when" always has an answer with a timestamp on it.

What happens if SystemsWarden itself goes down?

Nothing moves. Remediation is opt-in and fail-closed: with no confirmed multi-region verdict, no action is taken, and your existing DNS and proxy configuration keeps serving exactly as it was. Detection runs from edge locations independent of the control plane, so a control plane outage cannot invent a false failure either.

When can I actually buy this?

Not yet, and we would rather say so than take your card. SystemsWarden is in early access: accounts are free, billing is not switched on, and the prices on this page show where each plan is heading rather than what you will be charged today. Early-access users keep the limits we agree when billing does open.

Which alert channels are supported?

Email and Telegram today, with webhooks, Slack, Discord and SMS on the roadmap. Delays and thresholds are yours to set, so a two-minute blip on a staging box does not have to reach your phone.

Can I show uptime to my own customers?

Fleet includes white-label status pages, so your customers see your brand rather than ours, along with the uptime history that backs up whatever you promised them.

Get in while it is still being built.

Early-access accounts are free, keep the plan limits we agree, and shape what gets built next. Tell us what you run and we will get you set up.