Skip to main content

Harry detects stale Oracle performance ingestion before you get a call at 2am

True story: while updating Harry to v0.3.0, Grafana suddenly started /crying with a couple of nice firing alerts.

I was not testing alerting. Which made it even better. πŸ˜†

The accidental test​

While preparing Harry's version 0.3.0, I renamed the source databases in the scraper configuration. It was just a configuration update in the demo environment.

On the next dashboard refresh, Grafana started firing alerts!!

The previous source database labels had stopped receiving ingestion updates, and Harry's health checks correctly identified that their accounting and scraper data had become stale.

In other words: the system noticed that the data pipeline had stopped advancing for those sources.

What happened​

After the rename, the new source database names started to appear normally, while the old ones remained present in the repository without receiving fresh data.

Harry's alerting detected two different but related symptoms:

  • Harry repository ingestion accounting is stale
  • Oracle scraper data is stale
Active Alerts

This is useful because those alerts cover different layers of the monitoring path:

Oracle database
↓
Harry scraper
↓
PostgreSQL repository
↓
Grafana dashboards and alerts

If a dashboard still shows old data, but no new samples are arriving, that is dangerous. A stale dashboard can look calm and healthy while actually hiding a broken ingestion path. That is why stale-data detection matters.

Repository ingestion accounting alerts​

Harry keeps track of repository ingestion activity so it can detect when expected flushes are no longer happening.

In this case, the alert fired with the message:

No ingestion accounting flush has completed for more than fifteen minutes. Verify Harry leadership, PostgreSQL writes, and scraper logs. Accounting normally flushes every five minutes.

That alert is not only saying "something is wrong", it also points directly to the areas worth checking first:

  • Harry leadership
  • PostgreSQL writes
  • scraper logs
Harry repository ingestion Alerts

Oracle scraper data stale alerts​

At the same time, Grafana also raised an alert showing that Oracle scraper data itself had become stale for the old source database labels.

This gives a second confirmation that the data pipeline is no longer advancing as expected.

Together, these alerts help distinguish between:

  • a dashboard issue
  • a repository issue
  • a collection issue
  • or a broader scraper/runtime problem

Why this matters​

For me, this is one of the most important parts of observability: the monitoring system itself must also be monitored. It is not enough to monitor Oracle database performance if the monitoring pipeline itself can silently stop, slow down, or freeze.

I do not want a beautiful green dashboard if the data behind it stopped moving twenty minutes ago, only to get a call at 2 AM because something is really wrong.

Harry is designed to reduce that risk by exposing operational health and alerting on stale ingestion behaviour.

Harry 0.3.0 and self-monitoring​

This fits perfectly with Harry 0.3.0, which introduces Scraper pressure self-monitoring.

The goal is simple: Harry should not only collect Oracle performance data, but also provide visibility into its own runtime behaviour.

That includes detecting situations such as:

  • stale scraper output
  • delayed repository ingestion
  • ingestion surges or drops
  • runtime pressure conditions that may affect collection quality

This makes Harry more trustworthy in real operational environments, because the system can tell you not only what Oracle is doing, but also whether Harry itself is collecting and storing data in a healthy way.

Scraper Runtime Pressure on Harry Health dashboard Oracle Collector Duration on Oracle Operational Overview dashboard

Demo improvements​

As a bonus, this accidental test exposed one issue in the public demo: alerting was visible from the dashboard, but the read-only user could not inspect Grafana Alerting itself. That is fixed now.

See it in action​

You can explore Harry here:

Website: https://harryperformance.com

Demo: https://demo.harryperformance.com

This is a perfect example of how systems work in the real world: one small improvement can trigger a chain of events. Sometimes harmless, sometimes much more dangerous for the business. And that is exactly why monitoring the monitoring matters.