Monitor, investigate and recover Windows infrastructure from one place.
PulseWatch monitors servers, IIS, Windows Services, RabbitMQ and application health. It detects incidents, centralises the evidence and can safely run approved recovery actions — before your users notice.
- Queue depth rising on PROD-APP-03orders.process · 1,240 msgs, consumer rate falling
- RabbitMQ channel errors detectedPRECONDITION_FAILED · consumer ack timeout ×3
- Incident #1284 created — Criticalconsumer stopped processing · evidence captured
- Self-healing: service restartedRabbitMQ Consumer Service · rule “consumer stall”
- Queue back to baseline — resolvedMTTR 8 min · no human paged
Your whole Windows estate on one screen.
Every server, service, app pool and queue — health, incidents and self-healing activity, updating live.
Built for the stack you actually run.
See everything in one place
Servers, IIS app pools, Windows Services, RabbitMQ and disk/CPU/memory — one dashboard for the whole estate, refreshed by a lightweight agent every few minutes.
Incidents with the evidence attached
When something breaks, PulseWatch captures the moment: running services, error logs, processes and metrics at trigger time. No more "it was fine when I looked".
Recovery you can trust
Approved actions only — restart a service, recycle an app pool — with cooldowns, rate limits, dry-run mode and maintenance windows. Every action is signed and audited.
The consumer died quietly. Nobody got paged.
A RabbitMQ consumer stopped processing while its Windows service kept running — the failure mode that slips past service-level monitoring and gets discovered hours later as a 40,000-message backlog.
PulseWatch saw the queue growing, matched the channel errors, created the incident with the evidence attached, and restarted the consumer service under an approved self-healing rule. Twelve minutes, start to finish, fully audited.
- Queue depth starts increasing on orders.process — consumer rate falling below baseline.
- Consumer rate drops to zero. The Windows service still reports Running.
- Repeated RabbitMQ channel errors detected — PRECONDITION_FAILED, consumer ack timeout.
- Incident created with the full snapshot: services, logs, processes, metrics.
- Service restarted automatically self-healing under the approved “consumer stall” rule.
- Consumer restored — processing resumes, backlog draining.
- Queue returns to normal. Incident auto-resolved resolved
The 3am write-up, without the 3am.
Every incident carries its evidence — the services, logs, processes and metrics captured at trigger time. PulseWatch's AI reads that evidence and writes the analysis a senior engineer would: root cause, supporting evidence, what was done, and how to stop it happening again.
Bring your own provider — OpenAI, Anthropic or GitHub Copilot — and your incident data stays in your account. Analysis lands on the incident record, next to the timeline and the audit trail.
- Root cause (high confidence)
- Consumer channel closed by the broker — ack timed out while the thread blocked on a downstream call. The service host stayed alive, so service-level monitoring saw “Running” while the queue climbed to 4,180.
- Evidence
- 2× channel errors in the 60s before the incident · consumer rate 0/min while publish rate held · CPU/memory/disk nominal, ruling out resource exhaustion.
- Action taken
- Self-healing restarted the consumer service; backlog drained in 6 minutes.
- Prevent recurrence
- Set a client-side ack timeout · watchdog on consumer rate vs queue depth · add the missing HTTP timeout in the handler.
Run it on your own servers. Free for 30 days.
We're looking for a small number of teams running Windows infrastructure — IIS, Windows Services, RabbitMQ — to pilot PulseWatch on real workloads.
- Guided setup — the agent is a single PowerShell script
- No inbound firewall changes; agents push over HTTPS
- Self-healing starts in dry-run mode — nothing runs without your approval
- Direct line to the team building it
Or write to [email protected] — we reply within one business day.