Introduction
We are excited to introduce a regular part of the Operate Newsletter which is product enhancements. At our monthly reviews with all our Operate customers we discuss what enhancements we could make to our service that would improve their experience. Each month we aim to publish (and re-publish) the top 3-4 requests. In turn we ask you, the customer, to vote for the top enhancement which we will commit to incorporating into Apto Operate. So please grab a coffee or tea and have a read below voting in the poll at the end.
1. Skipped Search details
Today, Operate tells customers that a search was skipped, how many times it was skipped, and the general reason why. This gives some context, but in more distributed environments it doesn’t provide enough detail to rapidly identify the actual source of the problem.
A gap we could close is we don’t tell customers which specific component was affected. As an improvement, skipped search reporting should identify the exact search head where the issue occurred. This will work in both clustered, distributed and multisite environments, helping to identify unhealthy search heads, and highlighting underlying health issues. This would immediately narrow the scope of investigation and make the information far more actionable.
In addition, we should include the time the skip occurred. Many customers only retain search results for a short period, sometimes as little as one hour. When we say a search was skipped “at some point,” customers may check recent data and see no skips, which creates confusion. The issue isn’t trust; it’s a lack of precise timing. With timestamps, customers can correlate skips with known spikes in activity, competing workloads, and search schedules.
Providing both the search head and the exact time of the skip benefits everyone. Customers can diagnose issues more quickly, and our own investigations become more precise. This allows us to give a clearer, more complete explanation of why the search was skipped and what steps can be taken to prevent it in the future.
Instead of reporting, “This search was skipped four times due to too many concurrent searches,” we could say:
“This search was skipped four times on search head X between 10:15 and 10:30 due to high concurrent search load.”
This level of detail would significantly improve the value of skipped search reporting for Operate customers.
2. Source Type / Index Variance
This enhancement would monitor expected ingestion volumes across source types, indexes, or both. The goal is to establish what “typical” ingestion looks like over a defined period, then alert when volumes deviate beyond an agreed threshold, for example a 20% increase or decrease.
Defining “typical” and setting the right variance threshold will need careful implementation. A simple static threshold won’t work for all customers. For example, customers that batch ingest data would regularly trigger false alerts if we didn’t account for their normal ingestion patterns. To avoid this, we’d need an initial learning or analysis period per customer, per source type or index, to understand their natural variance and set thresholds accordingly.
Ideally, we would pilot this with one or two customers first and use that data to refine how thresholds are calculated before rolling it out more broadly. Over time, collecting ingestion patterns across customers would allow us to tune these variance models and make the alerts more accurate.
The benefit to customers is twofold. First, it provides early warning when ingestion suddenly spikes. For example, if a network device is left in debug mode and ingestion jumps by hundreds of gigabytes per day, the alert could flag that a specific index or source type is suddenly ingesting 200% more data than normal. This gives the customer a chance to investigate and disable the source before it impacts licensing or performance.
Second, it also detects unexpected drops in ingestion. If Windows event data falls by 25%, instead of a generic alert, we could indicate that roughly a quarter of Windows hosts have stopped reporting. This moves Operate beyond basic ingestion cessation alerts and towards more meaningful, actionable insights about data health.
Overall, this approach would make ingestion monitoring more proactive, more accurate, and far more useful for customers.
3. Ingestion Delay
This improvement would focus on measuring the delay between when an event occurs and when it becomes visible in Splunk. This time delta is critical, as it directly affects how quickly customers can detect and respond to issues.
For example, if an event occurs at 3:00 p.m. but Splunk doesn’t receive it until 3:30 p.m., there is a 30-minute ingestion delay. In a security scenario, such as an attempted breach, that delay is significant, and renders the common 1-10-60 rule impossible. By the time the event is visible, the incident may already be well underway or complete. Knowing about these delays in near real time is essential to maintaining a strong security posture.
The same issue applies to IT operations use cases. If telemetry from infrastructure is delayed, critical thresholds may be exceeded for extended periods without visibility. When the data finally arrives an hour later, the service may have already failed, removing any opportunity for proactive intervention.
By actively monitoring and alerting on event latency, Operate could surface when data is arriving later than expected and identify which sources or pipelines are responsible. This would give customers the ability to address delays before they impactsecurity or service reliability.
The main consideration with this approach is that it would require more frequent telemetry collection to ensure accuracy and reliability. However, the added visibility would significantly improve both operational awareness and response times for customers.
4. Roles Based Utilisation
This improvement would focus on monitoring usage at the Splunk role level to identify when users are nearing or exceeding their allocated utilisation thresholds/limits.
In Splunk, role limits such as search concurrency (the number of searches that can run at the same time) and disk usage for stored search results. Exceeding these limits should be blocked however we commonly see issue caused by role inheritance, or users with multiple roleswhich often leads to degraded performance and downstream issues. While Splunk provides default values out of the box, many customers create custom roles to support different teams or usage patterns.
The proposed enhancement would continuously monitor role utilisation and highlight when a role is consistently close to, or exceeding, its limits. This gives customers clear visibility into how capacity is actually being used across teams.
The benefit to customers is twofold. First, it enables informed tuning of role limits. If one user group is regularly hitting its limits, customers can justify increasing those thresholds to improve the user experience. Conversely, if another group is consistently well below its allocation, their limits could be reduced and that capacity reallocated where it’s needed most.
Second, this helps uncover hidden or unexpected heavy usage. We frequently see customers who own and pay for Splunk but are unaware of the full breadth of their user base, including which users or teams are driving the highest load. In many cases, a small subset of users is responsible for disproportionate resource consumption, and this goes unnoticed until performance issues arise.
This type of overutilisation is also one of the primary causes of skipped searches, which remains one of the most common alerts we see. By identifying and addressing role-level pressure early, Operate could help customers reduce skipped searches, improve platform stability, and deliver a more consistent experience across their Splunk environment.
See how we can build your digital capability,
call us on +44(0)845 226 3351 or send us an email…
-
1 July 2026
Building a Foundation for Risk
-
22 June 2026
The Convergence Imperative
-
4 June 2026
From Reactive to Resilient: Managed Splunk Operations for a Leading UK Financial Business


