Changelog

Follow up on the latest improvements and updates.

RSS

  1. Fixed an Overview panel crash.
    The Overview chart assumed every response would always contain at least two data series. When an application had zero or only one series, such as a new app with an empty state or a single metric, the panel tried to read a series that didn't exist and crashed instead of rendering. Overview now handles any number of series safely, from zero up, and each series gets its own color automatically instead of relying on fixed color slots.
  2. AWS accounts can always be removed cleanly, even after permissions change on the AWS side.
    Removing an AWS account now cleans up everything KloudMate stored for it directly, without depending on the AWS-side permissions still being active. If the connection is broken, you'll see a clear status banner explaining why, but delete stays fully working either way.
  3. Stability and display improvements across dashboards
    , including new
    Sections
    to group related panels,
    Repeating panels
    to generate one panel per value of a variable, a
    graph option inside Stat panels
    so a single number can show its trend alongside it, and
    variable values scoped to an attribute filtered by metric
    , so a dashboard variable only lists values that are actually relevant to the metric it's tied to.
Protection from day one:
new hosts get built-in CPU and memory safeguards that alert immediately, no baseline-learning wait.
Threshold
detectors need no history and fire the moment a value crosses a fixed, known-bad line, like a disk at 95% full or a container at 95% of its memory limit. While a brand-new host or pod is still building the history an anomaly detector needs, these threshold nets are already watching, so genuine saturation on a fresh entity is caught immediately instead of waiting for a baseline to form.
Fewer false alarms:
tuned anomaly detection so alerts are the ones worth acting on.
Anomaly
detectors learn what normal looks like for each entity from its own recent history and fire only when a reading lands well outside that range, with a
Sensitivity
setting (High, Medium, or Low) to control how tightly that range is drawn. A breach also has to hold for about 30 minutes across several evaluations before it fires, so a brief spike settles on its own instead of paging anyone.
smart-alerts-detectors
Check out the Smart Alerts guide for the full breakdown.
  • Separate iOS app hangs from Android ANRs to diagnose platform-specific issues.
    The Errors tab now splits by kind:
    All errors
    ,
    Crashes
    , and
    App Hangs
    , with
    ANRs
    shown in place of App Hangs when you're looking at an Android app. The same distinction carries into a session: the header badges show crashes, app hangs, and ANRs separately, and they appear as their own markers on the replay timeline. On iOS specifically, an app-hang stack is sampled just before the main thread blocks, so it points near the cause of the hang rather than exactly at it.
  • Track slow and frozen frames per screen to find where your app feels sluggish.
    Slow and frozen frames are captured automatically once the SDK initializes, alongside app start, device vitals, and tap spans. On the Performance tab, Android and iOS apps get a dedicated
    Rendering
    section, alongside App start, Screen loads, and Who is slow, so you can see which screens are dropping frames instead of guessing from crash reports alone.
  • A dedicated Errors view to search, filter, and understand what's breaking.
    Errors now has its own facet rail for narrowing the list, with a volume chart above it.
    Grouped by issue
    collapses identical errors into a single row with its count and first and last occurrence, and
    Raw events
    switches to individual occurrences when you're chasing one specific failure rather than a pattern. Opening any error shows its full stack trace, its attributes, and every session it happened in.
Check out the RUM Interface guide and Session Detail guide for the full details.
Understand how users move through your product.
  • Conversion Funnels:
    build funnels to see where users drop off, with drop-off reasons that explain why, plus smart funnel suggestions to get started faster. A funnel measures conversion across up to five steps, where each step is either a route (like
    /cart
    ) or a custom event (like
    checkout_completed
    ), and either can match a prefix. Above the chart you get
    Conversion
    ,
    Biggest drop
    , and
    Time to convert
    , and each step reports its own session count and drop-off.
    Complete within
    controls how long a session has to move from the first step to the last before it stops counting as a conversion, and
    Compare to previous period
    shows the change in percentage points against the prior window. Clicking any step opens the sessions that stalled there, along with the errors they hit.
    Save a funnel
    to track its conversion over time.
rum-journeys-funnels
  • On-demand history:
    backfill a saved funnel's metrics so you can analyze past performance immediately, not just from today forward.
  • User Journeys:
    follow the real paths users take across your app.
    Pathways
    draws the routes people actually took as a left-to-right graph, with each node showing a page's session count and exit rate, and each edge showing how many sessions took that path. Clicking a node redraws the graph starting from that point and opens the sessions behind it, along with the share who
    left there
    or
    hit an error
    , so a high exit rate mid-path is easy to spot.
rum-journeys-pathways
  • Custom Events:
    track the product events that matter to your business and use them across RUM. Anything sent through
    addEvent()
    shows up in
    Analyze
    , where you pick a measure, an aggregation, and a group-by to render it as a table, a trend, a top list, or a distribution. Custom events are counted per session and per user, and can be used directly as funnel steps.
rum-events-over-time
Check out the RUM Interface guide for the full breakdown.
See exactly how people experience your apps, on web and mobile.
Session Replay:
play back real user sessions with a smoother player, hour-long session support, and "skip inactivity" to reach the moments that matter. Playback runs from 1x up to 8x speed, and Skip inactivity jumps past the quiet stretches that make up most of a long session. Replay is tied to the same timeline as the rest of the session: clicking an error or a frustration click in the event list moves the replay straight to that moment.
Mobile RUM:
new support for native iOS, React Native, and Expo apps, with the same insights you get on the web. Mobile apps now share the SDK's collector and the same RUM interface as your website, so a workspace can hold a website and its mobile apps side by side. Screen views, app start, crashes, and tap spans are captured on iOS and React Native the same way they already are on Android, with a guided setup for each platform.
Rage & dead click detection:
accurately spot where users get stuck or frustrated. A rage click is three or more clicks on the same element within one second. A dead click is a click that produced no DOM change and no navigation within 500ms. An error click is a click followed by a JavaScript error within one second. All three show up as badges on the sessions list, as markers on the replay timeline, and as rows in the session's event list, so you can jump straight to the click that frustrated someone.
Redesigned Performance view:
clearer charts and readable tables for how your pages and endpoints really perform. Performance now opens on Core Web Vitals at the 75th percentile, with a distribution view underneath that shows whether a handful of visitors had a bad time or everyone had a mediocre one. Slowest pages by impact ranks pages by total load time contributed rather than by the raw measurement alone, and an endpoints table puts latency and failure rate on the same row.
Guided setup:
a step-by-step flow to add a new application in minutes. The Add application wizard walks through picking a platform (web, Android, or iOS) and generates a ready-to-use snippet with your key and current SDK version already filled in.
Check out the Instrumentation Guide for full setup instructions.
Alert rules can now catch an instance that stops reporting entirely, not just one that crosses a threshold. Turn it on per rule, and KloudMate remembers every instance it's seen, then alerts if any of them goes silent.
An instance is a unique series identified by its labels; group a heartbeat by host_name, and you get one instance per host; group by serviceName + pod_name and you get one per pod. Each is tracked individually, with its own state, history, row on the Instances tab, and dismissal. It's off by default and opt-in per rule, so nothing changes for existing rules until you turn it on. It works with any datasource.
How it fires: an instance missing from query results moves through the normal alert lifecycle, Pending, then Firing, with the reason instance stopped reporting; last seen <timestamp>. It recovers automatically the moment it reports again. Roughly, it fires the query window plus the pending duration after the last data point, so the window should be several times your reporting interval; a 60-second heartbeat with a 60-second window will flap on one late data point; a 5-minute window absorbs the jitter.
Cleanup that doesn't require babysitting:
  • Dismiss an instance you decommissioned on purpose; one click, closes right away
  • Auto-close handles the rest automatically, once the configurable window passes (default 24h) while other instances are still reporting
  • If the whole query goes dark instead, nothing auto-closes; every tracked instance keeps firing until data returns or you dismiss it, since a totally dark query could just mean a broken pipeline, not a dead fleet
Check out the Instance Absence Detection documentation for full details, including how it interacts with routing rules, folders, silences, and maintenance windows.
Kubernetes Monitoring now includes a Storage tab showing Persistent Volume Claims across your clusters, giving visibility into storage alongside your existing pod and node metrics.
Each row shows the Persistent Volume name, cluster, status, claim name, storage class, and capacity. Filter by Persistent Volume type, cluster, or status, or search by name.
Screenshot 2026-08-10 210639
This should make storage issues visible in the same place you already check pod and node health, instead of needing kubectl to see what's bound, how much capacity is used, or what's stuck pending.
Check out the Kubernetes Monitoring documentation for full details.
Sharing a dashboard now works the same way no matter where you start it; sharing from the dashboard list and sharing from an individual dashboard page both open the same Share dialog with both options available, instead of the list forcing an immediate JSON copy while the dashboard page only offered a public link.
  • Public link:
    generates a read-only link viewable without logging in, with toggles for enabling it and letting viewers change the time range, managed afterward from the Public Dashboards page
  • Copy JSON:
    copies the dashboard's full definition; paste it into another workspace's Import dialog to recreate the same dashboard with no manual rebuild
Screenshot 2026-08-10 194805
Screenshot 2026-08-10 194457
Check out the Dashboard Details documentation for how sharing works.
Correlates related alerts into a single incident instead of flooding the team with separate alarms. It's opt-in per routing rule: switch a rule's Alert grouping setting to
Auto (AI)
, and the engine takes over from there, learning which alerts tend to fire together over time. No grouping keys to set.
When multiple alerts fire close together, the engine works down a signal ladder to decide whether, and why, they belong together, checked in order:
  • Same alarm rule
    firing on multiple targets, always merges. A rule firing across many hosts becomes one group with many instances, not a separate ticket per host
  • A shared identity label;
    both key and value must match
    (host_name=fedora
    and
    host_name=ubuntu
    do not match)
  • Co-fire history
    , alerts that have repeatedly fired together before
  • A known cause-effect pattern
    , e.g. a DynamoDB throttle causing Lambda errors
  • Topology
    , the entity graph connects the two resources
  • An LLM pass
    , for novel combinations nothing above can explain
Each group shows its members and a
"Why grouped"
reason. You can thumbs-up or thumbs-down a grouping decision; a single thumbs-down is a one-off override, but a second person flagging the same pairing stops it from grouping that way again. Groups link directly into
Run RCA
for investigation.
Screenshot 2026-07-31 at 11
1. AWS dashboard panel edit mode not populating data source/query with custom data source variable.
When a panel was built using a custom data source variable, reopening it for editing showed the data source correctly selected but left the query field empty, even though a valid query had been saved. This made a working panel look unconfigured and risked someone overwriting it with a blank query on save. Both the data source and the previously saved query now reload correctly in edit mode.
2. AWS alert dimension labels merging together in query preview with multiple dimensions.
When building an alert query with two or more dimensions (e.g.,
service
+
error_type
), the preview concatenated the labels into one unreadable string instead of showing them separately, also affecting Dashboards and the Explore page. This made it impossible to verify a multi-dimension query was configured correctly before saving. Labels now display cleanly and separately per dimension.
3. Report Download button failing due to CORS error.
Clicking Download on any report failed silently, with a blocked CORS request in the browser network log, meaning no reports could be exported from the UI at all (emailing the report was the only workaround). The cross-origin request issue is fixed, and downloads now complete normally.
Load More
→