Back to Studio

AI-friendly Markdown · structured for AI citations

Growth · Studio

How to Monitor a Side Project During Your Day Job

14 min read
How to Monitor a Side Project During Your Day Job

Monitor a side project during your day job: uptime checks, error alerts and job heartbeats, split so only real outages reach your phone and the rest waits.

Here's how most side projects find out they're broken. You're in a meeting at your day job. Your signup form stopped working at 9:40 that morning. Nobody tells you. At 8pm you open your laptop, see zero new users for the day, and only then notice the error. Or worse, a stranger mentions it in a reply you read two days later.

You can't watch a dashboard during work hours, and you shouldn't try. Monitoring a side project isn't about staring at graphs. It's about deciding which few failures are allowed to interrupt your day, sending those to your phone, and letting everything else wait until the evening. Set it up once on a weekend and it mostly runs itself.

Below: what to watch, how to split alerts into "now" and "tonight," the three checks that cover most small products, and a ten-minute phone triage. Tool plans and limits were checked on each vendor's own page in October 2026, so check again before you rely on one.

What should you actually monitor on a side project?

Monitor the few things that hurt users when they break and that you wouldn't notice on your own: the site responds, the core action works, scheduled jobs run, and email goes out.

A useful idea from Google's SRE chapter on monitoring distributed systems is the difference between symptoms and causes. A symptom is what the user feels ("signup returns an error"); a cause is why ("the database refused connections"). The chapter recommends putting much more effort into catching symptoms. Good news for one person: symptoms are easier to check from outside, and there are fewer of them.

For each part of your product, finish the sentence: "If this broke at 10am on a Tuesday, I'd find out when..." If the honest answer is "when a user complains," it goes on the list. For a typical side project:

  • The app is reachable. Not only the marketing page, but the part people log into.

  • The core action works. Signup, login, and the one thing people came for, like creating a report or sending a reminder.

  • Scheduled jobs actually ran. Nightly backups, reminder emails, data imports, certificate renewals.

  • Outgoing email is sending. Password resets and verification emails, the ones users can't work around.

  • New errors in your code. Exceptions that didn't exist yesterday, usually from your latest deploy.

Which alerts should reach you at work, and which can wait until tonight?

Only alerts that mean users are blocked right now, and that you can act on from your phone, should interrupt your workday. Everything else belongs in an evening review.

The SRE chapter says every page should be actionable, and that people who get paged too often start skimming or ignoring alerts, sometimes including the real one. That happens to full-time engineers. It'll happen faster to you between meetings. So split alerts into two tiers before you set up any tool.

Tier 1: interrupt me now

  • The app or its health check is down for two checks in a row.

  • The core action fails (for example, your signup check returns an error).

  • A time-sensitive scheduled job didn't run, like reminders users expect at 5pm.

Aim for Tier 1 to fire a few times a month at most. More than that means something is mis-tuned or fragile, and either deserves a weekend fix.

Tier 2: show me tonight

  • New error types, and old errors that came back.

  • Slow responses that aren't outages.

  • A non-urgent job missed one run, like a weekly report.

  • Certificates or other things expiring in the next few weeks.

Tier 2 goes to email or a chat channel you check after work. The SRE chapter notes that email alerts easily get overrun with noise, which is why they shouldn't carry anything urgent. Email for the evening pile, push notifications for Tier 1.

How do you set up uptime checks that don't cry wolf?

Point an outside checker at a health endpoint that touches your database, have it look for a specific word in the response, and only alert after the check fails twice in a row.

Check from outside, and check the right URL

An uptime check hits your site from someone else's servers, the way a user would. The SRE chapter calls this black-box monitoring. The trap is checking your homepage: a cached or static homepage can load perfectly while the app behind it is dead. Check the app domain or an API route instead.

Add a small health endpoint

Create a route like /health that runs a trivial database query and returns 200 if it worked, 503 if it didn't. Keep it fast, skip authentication, and leave error details out of the response. That turns "the server answers" into "the server can do work."

Look for a word, not just a status code

The SRE chapter counts a 200 response with the wrong content as an error. Most uptime tools offer a keyword check that passes only if a specific word appears. Pick one that only shows up when things worked, like "ok" in your health response.

Pick a sensible interval and a confirmation rule

The SRE chapter suggests that even a service targeting 99.9% uptime probably doesn't need to probe more than once or twice a minute. For a side project, five minutes is fine. UptimeRobot's free plan, which it pitches at hobby and non-profit projects, includes 50 monitors at a five-minute interval and keyword monitoring (as of October 2026). Then turn on your tool's setting to wait for a second failed check, or a few minutes of downtime, before alerting.

How do you catch errors an uptime check can't see?

Add an error tracking tool to your app so exceptions get recorded with a stack trace, and alert on new kinds of errors rather than every single one. An uptime check tells you the site is up. Error tracking tells you whether it's working for the person who just clicked "save."

You install a small library, it catches unhandled exceptions in your backend and frontend, and it groups repeats into one "issue." Sentry's Developer plan is free for one user and includes error monitoring with alerts and notifications by email (as of October 2026). Email alerts fit Tier 2 nicely.

The useful part is choosing triggers. In Sentry's guide to creating alerts, the triggers include "a new issue is created" and "a resolved issue regresses," and you can filter by environment and throttle how often an alert repeats for the same issue. For a side project, a good default is:

  • Alert when a new issue appears in production.

  • Alert when a resolved issue comes back, because that usually means a deploy undid your fix.

  • Throttle repeats to once a day per issue, so one bug doesn't send fifty emails.

  • Ignore staging and your local machine completely.

In the first week, spend twenty minutes on noise. Browser extensions, ad blockers and bots throw errors you can't fix, so ignore them once you've confirmed they aren't yours. Upload source maps if your frontend is bundled, so stack traces point at real code. And tag each deploy with a release name, so you can see which deploy brought a new error in.

How do you know a scheduled job silently stopped?

Use a heartbeat check: the job pings a unique URL every time it finishes, and the monitoring service alerts you when a ping doesn't arrive on time. It catches what every other check misses: something that should have happened and didn't.

Scheduled jobs fail quietly. The scheduler doesn't come back after a restart, or an environment variable changes and the job exits early. Nothing errors, because nothing ran. The Healthchecks.io documentation describes this pattern as a dead man's switch: the service keeps silent as long as pings arrive on time and raises an alert as soon as one doesn't.

Each check has an expected period and a grace time. The docs' example: for a job that should run every hour with five minutes of grace, if the last ping arrived at 12:00, the check is marked late at 13:00 and the alerts go out at 13:05. You can also send an explicit failure signal by adding /fail to the ping URL, so a job that crashes tells you right away instead of waiting out the grace period. Healthchecks.io's free Hobbyist plan monitors up to 20 jobs (as of October 2026).

Put a heartbeat on these first:

Two details matter. Only ping on success (in a shell script, chain the ping after your job with &&). And treat ping URLs as secrets, as the docs recommend, because anyone who has one can send fake signals. Keep them out of public repos.

How do you get urgent alerts through on your phone without wrecking your workday?

Send Tier 1 alerts to one push notification app, allow only that app through your phone's focus or do-not-disturb mode during work hours, and send everything else to email. Then add a second channel for Tier 1 so one failure doesn't leave you blind.

On iPhone, Apple's guide to allowing notifications during a Focus shows how to pick apps that can still reach you, for example in a Work Focus. On Android, Google's guide to Modes and Do Not Disturb lets you choose which apps can notify you under each mode's notification filters. Let the monitoring app through and keep project email out.

Healthchecks.io's notes on configuring notifications suggest two different channels, so if one fails (an email lands in spam, say) the other still reaches you, and different methods depending on urgency, which is the two-tier split. The same page mentions Pushover's Emergency priority, which plays a loud sound every five minutes until you acknowledge it. Most side projects don't need that.

Also decide what happens at night. If you won't get out of bed for your side project, and most people shouldn't, route nothing to your phone overnight. Be honest about that in public and don't promise response times you can't keep from a day job. If you sell to teams, the advice on writing a security page that unblocks your first B2B deals covers how to describe your incident process honestly at one-person scale.

What should you do when an alert fires in the middle of the workday?

Follow a ten-minute phone triage you wrote down in advance: confirm it's real, undo your last deploy if that's the cause, tell users in one line, and leave the real fix for the evening.

  1. Confirm. Open the site on your phone over mobile data, not office Wi-Fi.

  2. Ask "did I deploy?" Breakages often follow a deploy. If you shipped recently, roll back first and investigate later. On Vercel, for example, Instant Rollback lets Hobby users roll back to the immediately previous deployment. One gotcha from those docs: after a rollback, new pushes to your production branch won't go live automatically until you undo the rollback. Note that in your runbook.

  3. Check your providers. Your host, database or email provider might be the one that's down.

  4. Tell users in one line. A short "we know, it's being fixed" stops people guessing. If you don't have a place for that yet, here's why it's worth it to ship a status page before your first real outage.

  5. Stop at ten minutes. If rollback didn't help and no provider is down, accept a few hours of downtime and fix it properly tonight.

Keep it all in one phone note: dashboard links, rollback steps, your status page login and your providers' status pages.

How do you test that your alerts actually work?

Break things on purpose once, on a weekend, and watch each alert arrive with your work focus mode on. An alert you've never seen fire is a guess, not a safety net.

  • Uptime: make /health return 503 for ten minutes. You should get one alert, then a recovery message.

  • Errors: add a test route that throws an exception in production, hit it once, then delete it. Sentry's alert builder also has a "Send Test Notification" button to check the wiring.

  • Heartbeat: skip one run of a scheduled job, or point it at the /fail URL once.

  • Phone: if an alert doesn't get through your work mode, fix the settings now, not during a real outage.

Repeat the phone test whenever you get a new phone or change notification settings.

What does a 15-minute weekly alert review look like?

Once a week, list every alert that fired, and ask three questions about each. Then fix, re-tier or delete it. Fifteen minutes keeps the system quiet enough to trust.

  1. Was it real? If not, tighten it: add a confirmation rule, a keyword, or an ignore filter.

  2. Did I act on it? If it was real but you didn't need to do anything, move it to Tier 2.

  3. Did anything break that no alert caught? If a user reported something first, add the missing check.

The SRE chapter treats rarely exercised alerting rules, and signals nobody looks at, as candidates for removal. Same idea here: if an alert hasn't helped in three months, question it. This review should fit inside the small weekly ops budget you set when scoping a side project MVP you can ship on weekends. If monitoring keeps eating more, the product needs a stability weekend.

What does this look like on a small side project?

Here's an example. Everything in it is made up: the person, the product and the numbers.

Ana works full time as a data analyst. On evenings and weekends she runs Plotbell, a small web app that sends watering-rota reminders to members of community gardens. Every day at 5pm a scheduled job emails whoever is on watering duty. About 300 gardeners across 12 gardens rely on it.

Her setup took one Saturday afternoon:

  • One uptime check on /health every five minutes, with a keyword match and an alert after two failures, sent to a push app.

  • One heartbeat on the 5pm reminder job, with a 30-minute grace time, sent to the same push app and to email.

  • One heartbeat on the nightly backup, sent to email only (Tier 2).

  • Error tracking with alerts for new and regressed issues, by email, throttled to once a day per issue.

  • A Work Focus on her phone that lets only the push app through.

In week three, a deploy she pushed the night before renamed an environment variable, and the reminder job exited without sending anything. No error, no outage. At 5:30pm the heartbeat went past its grace time and her phone buzzed on the train home. She opened her runbook note, rolled back to the previous deploy from her host's dashboard, triggered the job by hand, and the reminders went out at 5:50. She fixed the variable properly that evening.

Without the heartbeat, she'd have heard about it the next morning, from a gardener asking why nobody watered the tomatoes.

What are the most common monitoring traps on side projects?

  • Watching the homepage only. The marketing page stays up while the app is broken.

  • Alerting on every error. You'll mute the whole thing within a week.

  • One channel only. One spam filter or expired app login and you hear nothing.

  • Never testing. The first real alert is the worst time to learn your phone silences it.

Quick answers

Do I need monitoring before I have real users?

A minimal version, yes: one uptime check and one heartbeat on your backups. Add the rest once people rely on the product.

How often should uptime checks run for a side project?

Every five minutes is plenty. Pair it with a two-failure rule so you only hear about real outages.

A side project should be able to break without you finding out a day late. Pick your five checks, split alerts into "now" and "tonight," test them once with your phone in work mode, and review them weekly. When your project is steady enough to show more people, SideHunt is a friendly place to launch it and find its first users.

side projectsmonitoringuptimeerror trackingalertsindie