Incident management that doesn't assume you have an incident commander.
A timeline per incident, a status page that updates itself, and acknowledgements that work from your inbox at 11pm. Built for teams where the person on call is also the one answering tickets.
No war room to staff. No second tool to wire up to your status page. The incident is the status update.
API latency elevated for some customers
InvestigatingTimeline
We're seeing elevated API latency for some customers and are looking into it. We'll post an update here as we learn more.
Priya
Looks correlated with this morning's deploy — pulling the slow-query logs now.
Sam
You don't need a war room. You need fewer tabs.
Here's how an incident actually goes on a team your size. Something looks off; someone posts in Slack — "is the dashboard slow for anyone else?" An hour in, a customer emails. "Should we update the status page?" — but that's a different tool nobody has open, run by the same person reading the logs. There was no war room. There was you, a Slack thread, and three tabs at 11pm.
A timeline, not a war room.
An incident moves through four states. That's the whole flow — no phases to configure, no roles to assign.
Something's wrong and you're working out what. The default starting state.
You know the cause, and can tell customers — even before it's fixed.
The fix is in; you're watching to be sure. Doesn't count against uptime.
Done. Posting the resolving update stamps the resolved time.
Severity, and what your customers see.
Severity is yours to set internally — but it also decides, automatically, what customers see on the status page. Four levels, four public states, mapped one-to-one.
Two outage states, because "down for everyone" and "broken for some" aren't the same promise. You set the severity once; the public label follows.
Set the severity. Your status page already knows.
Most tools make you do this twice — manage the incident here, re-type it into a status-page tool there, kept in sync by hand mid-incident.
In StayUpfront the status page is a view of the incident, not a copy. Set the severity and pick the affected components; the public page reflects it the moment you post, and clears itself when you resolve. Your internal severity stays internal — customers see the plain-English version, and you decide what's public per update.
The status page is a derivative. Severity, component status, public label — automatically. No second tool, no manual paste.
Resolve once. Move the incident to resolved and the affected components clear themselves. No "remember to flip the status page back".
Public is a choice, per update. Each update can be public or internal. Customers see the wording you chose — not your internal severity label.
API latency elevated for some customers
MonitoringWe've rolled back the slow query and load times have recovered. We're watching to be sure before we resolve.
Updated
One incident. The left is what you set; the right is what your customers see. You only touch the left.
Incidents start themselves. Or you start them.
When a web check fails, a job misses its heartbeat, or a domain check spots an expiring certificate, StayUpfront opens the incident for you — linked to the monitor that raised it, resolved when it recovers.
Spot something a monitor can't? Declare it yourself — same incident, same timeline, same public surface.
Auto-created from monitors. Web checks, heartbeat monitors, and domain checks (which cover DNS and certificate expiry) open incidents on their own — tagged as automated, linked to their source.
Or declared by hand. Something a monitor can't see? Open an incident in a couple of fields. Same timeline, same public surface.
Heartbeat overdue: nightly-export job
InvestigatingResolves automatically when the monitor recovers.
One timeline. Internal notes and public updates, in order.
Every incident is a single timeline, posted in order. Public entries are the updates customers read on the status page; internal notes are marked with a lock and never shown publicly — same thread, so team context and the customer-facing story aren't in two places.
The timeline is the record. Change the incident — title, severity, status, components — and the change is logged automatically as an internal note. Timestamps stay in order and the first update can't be deleted. The history holds.
Timeline
We've rolled back the slow query and load times have recovered. We're watching to be sure before we resolve.
Priya
Slow query traced to this morning's deploy. Rollback running now — hold the public update until latency drops.
Sam
Severity changed from major to minor.
Priya acknowledged this incident.
Acknowledge from the email. At 11pm. Without opening anything.
The alert email has an acknowledge link — tap it from your phone, no login. (There's an in-app button too.) The acknowledgement lands on the timeline, so the team can see someone's got it.
Acknowledging stops the escalation — if a voice call was queued to wake you, it doesn't call.
Ack by email link. A tap on the alert email acknowledges the incident — no app, no login. There's an in-app button too.
Acknowledged means acknowledged. Escalation stops, and a queued voice call checks your ack status before it dials. Already on it? It won't call.
An incident on API needs an owner. Acknowledge to let the team know you're on it and stop the escalation.
AcknowledgeWorks straight from this email. No app to open.
The ticket that reported it, right there on the incident.
When a customer raises a ticket about the thing that's breaking, you shouldn't play detective across two tools to connect them.
In StayUpfront, incidents and tickets link both ways: the incident lists the tickets raised about it, and each ticket shows the active incident. Declare an incident from a ticket and it pre-fills the severity, visibility, and affected components. The agent sees "this is a known incident, here's the status" instead of replying "have you tried a hard refresh?"
Linked tickets
Dashboard slow to load this afternoon
Getting timeouts on the API since ~2pm
On-call, for when you grow into it.
You don't need a rotation — a founder and one engineer can run incidents with every alert going to both. When you want one, it's here: weekly or daily schedules, overrides for swaps and holidays, and a dashboard showing who's on call and where the gaps are. Alerts go to whoever's on call; with nobody covering, they fall back to the whole team rather than alerting no one.
Slack threads, both ways.
Incident updates post to one Slack thread; replies there flow back into the incident record.
Big enough? Still fits.
Most incident tools assume a platform team before you can run a clean incident. This doesn't — and the shape holds as one person grows into three, then a whole team: declare it, work it, tell your customers, resolve it.
You'll know if this is you.
You're a founder and one engineer. "Incident management" sounds like a thing bigger teams do — but you're the one who notices the site's slow, fixes it, and should be telling customers. The incident, the status update, and the customer who reported it are one workflow here.
You're a small team tired of paying for three tools. A status-page tool, an on-call tool, and a support tool that don't know each other. Here the monitor opens the incident, the incident updates the status page, and the ticket that reported it is right there in the sidebar. One workspace, one bill.
Either way: incident management sized for the team you actually have.
Want this before your next incident? Get in early.
Private beta is a few weeks out, and I'm letting people in a small group at a time so I can actually work with each of you — what's breaking, what's missing, what to ship next. The earliest customers shape how this handles a real incident, because they're the ones running real incidents through it.
Drop your emailDirect email from Rob when your slot's ready. No drip sequence.