• Wrong document? Swap it for free
  • Written by students who passed
  • Immediately available after payment
  • Read online or as PDF
Sell
Where do you study
Your language
Document preview thumbnail
Preview 4 out of 31 pages
Exam (elaborations)

RTPM Test Exam 2026 - 90 Questions and Correct Answers (Microservices, Biostatistics, Biochemistry)

Document preview thumbnail
Preview 4 out of 31 pages

This comprehensive resource provides 90 expert-verified questions and correct answers for the 2026 RTPM (Real-Time Performance Monitoring) test exam. The document covers a vast range of topics essential for success, making it an invaluable study tool for students and professionals in computer science, data science, and related fields.

Content preview

RTPM TEST EXAM AND CORRECT ANSWERS.


1. In the context of real-time performance monitoring (RTPM) for a distributed microservices
architecture, which metric is most indicative of a systemic bottleneck that cannot be resolved by
simply scaling individual service instances?

A. Elevated p99 latency across multiple services with low CPU utilization on all nodes
B. High request rate to a single service with correspondingly high error rate
C. Consistent 95th percentile response time increase during peak load that correlates with increased database
connection pool usage
D. Spike in memory utilization on a single service instance followed by automatic restart

Answer: A
Rationale: Elevated p99 latency with low CPU suggests a contention or coordination bottleneck (e.g.,
lock contention, network I/O) that horizontal scaling cannot fix. Option B points to a single-service issue
solvable by scaling; C suggests a database bottleneck that might be alleviated by connection pooling or
replication; D indicates a memory leak that can be fixed by patching the instance.


2. A company implements RTPM to detect anomalies in a high-frequency trading system. The
monitoring system uses a sliding window of 1 second and flags any deviation beyond 3 standard
deviations from the rolling mean. During a flash crash, the metric values change so rapidly that the
flagging system fails to trigger until after the crash has already occurred. Which fundamental
limitation of the monitoring approach is most likely responsible?

A. The sliding window is too long, causing the baseline to adapt too slowly
B. The standard deviation is too sensitive to outliers, causing the threshold to be too wide
C. The metric being monitored is not suitable for anomaly detection in financial contexts
D. The 3-sigma rule assumes a normal distribution, which is violated during extreme market events

Answer: D
Rationale: Financial metrics during flash crashes exhibit fat tails and non-normal distributions, so the
3-sigma rule fails to detect anomalies because it underestimates the probability of extreme values.
Option A is incorrect because a 1-second window is already very short; B is opposite—outliers inflate
sigma, making threshold wider, not narrower; C is plausible but less fundamental than the distributional
assumption.


3. In RTPM for a cloud-native application, you observe that the 'request latency' metric shows
periodic spikes every 5 minutes, coinciding with garbage collection (GC) pauses in a Java-based
service. The service runs on a Kubernetes cluster with CPU limits set to 2 cores. Which adjustment
would most likely reduce the frequency and impact of these GC pauses?

A. Increase the CPU limit to 4 cores to allow parallel GC to run faster
B. Reduce the heap size to force more frequent but shorter GC cycles
C. Switch to a concurrent garbage collector and tune the GC to be more CPU-intensive
D. Implement a circuit breaker to drop requests during GC pauses




Page 1

,Answer: C
Rationale: Concurrent GC (e.g., G1) runs in parallel with application threads, reducing pause times at
the cost of higher CPU usage. Given the CPU limit, tuning for concurrency can spread GC work across
the 2 cores, minimizing stop-the-world pauses. Option A might help but is a resource increase, not a
tuning; B would increase frequency and might worsen impact; D is a workaround, not a root cause fix.


4. A DevOps team uses RTPM to trigger auto-scaling of a web application based on a composite
metric that combines CPU utilization and request queue depth. The scaling policy is: if CPU >
80% OR queue depth > 100 for 2 consecutive minutes, add one instance; if CPU < 30% AND queue
depth < 10 for 5 consecutive minutes, remove one instance. During a traffic surge, the system scales
up correctly, but after the surge subsides, the system repeatedly scales down and then back up,
causing oscillation. Which modification to the policy would most effectively prevent this
oscillation?

A. Increase the cooldown period after scale-down actions
B. Change the scale-down condition to require CPU < 20% AND queue depth < 5 for 10 minutes
C. Add a hysteresis band such that scale-down requires CPU < 40% AND queue depth < 50
D. Use a more aggressive scale-up threshold to handle surges faster

Answer: C
Rationale: Oscillation occurs because the thresholds for scaling up and down are too close, causing the
system to flip-flop. Adding hysteresis (asymmetric thresholds) creates a dead zone where no scaling
occurs, stabilizing the system. Option A only delays the next scale-down but doesn't prevent oscillation;
B makes scale-down harder but still has a narrow band; D addresses scale-up but not the oscillation.


5. In an RTPM system that tracks service-level objectives (SLOs) using a burn-rate alerting
approach, which scenario would trigger an immediate critical alert if the SLO is 99.9% availability
over a 30-day window and the burn rate is measured over a 1-hour window?

A. Error rate of 5% sustained for 10 minutes
B. Error rate of 0.1% sustained for 2 hours
C. Error rate of 1% sustained for 1 hour
D. Error rate of 10% sustained for 5 minutes

Answer: A
Rationale: Burn-rate alerting calculates how fast the error budget is consumed. With a 99.9% SLO, the
error budget over 30 days is 0.1% of total requests. A 5% error rate for 10 minutes would consume a
large fraction of the budget quickly, triggering a critical alert. Option B (0.1% for 2 hours) is within
budget; C (1% for 1 hour) is less severe; D (10% for 5 minutes) is severe but shorter duration, but still
less total budget consumption than A.


6. A monitoring team implements distributed tracing across a microservices architecture using
RTPM. They notice that a particular trace always shows high latency in service A, but when they
drill down, the latency is due to service A waiting for a response from service B. However, service
B's own metrics show low latency. What is the most likely cause of this discrepancy?

A. Service B is experiencing a network delay that is not captured in its own metrics
B. Service A is using a synchronous blocking call that is queued at service B due to thread pool exhaustion



Page 2

,C. The trace sampling rate is too low, causing misleading data
D. Service B's latency measurement does not include the time spent in its own internal processing

Answer: B
Rationale: If service B's thread pool is exhausted, requests from A are queued at B before processing, so
B's own latency metrics (which measure processing time) appear low, but the total time from A's
perspective includes queue wait. Option A would affect B's metrics too; C would cause sampling errors
but not this pattern; D is opposite—B's latency should include processing.


7. In RTPM for a content delivery network (CDN), which metric would best indicate that a
particular edge node is experiencing a cache thrashing problem?
A. High cache hit ratio with low CPU usage
B. Low cache hit ratio with high disk I/O and eviction rate
C. High cache miss ratio with low network bandwidth
D. Low cache miss ratio with high memory usage

Answer: B
Rationale: Cache thrashing occurs when the working set exceeds cache capacity, causing frequent
evictions and reloads, leading to low hit ratio and high disk I/O. Option A indicates good caching; C
might indicate a different issue; D indicates efficient caching.


8. A company uses RTPM to monitor a data pipeline that processes streaming data with Apache
Kafka. The monitoring shows that the consumer lag for a particular topic is increasing steadily
over time, but the CPU and memory usage of the consumers are low. Which of the following is the
most likely root cause?

A. The number of partitions in the topic is too low for the desired throughput
B. The consumers are blocked on an external API call with a long timeout
C. The Kafka brokers are experiencing disk I/O bottlenecks
D. The message size is too large for the network

Answer: B
Rationale: Low CPU/memory but increasing lag suggests consumers are spending time waiting (e.g., on
I/O or network calls) rather than processing. Option A would cause high CPU if consumers are busy; C
would affect brokers, not consumers; D would cause network issues but not necessarily low CPU.


9. In RTPM for a machine learning inference service, which metric would be most appropriate for
detecting model drift in real-time?
A. Inference latency p99
B. Prediction confidence scores distribution
C. Number of requests per second
D. Memory utilization of the model server

Answer: B
Rationale: Model drift refers to changes in the input data distribution or the relationship between inputs
and outputs, which can be detected by monitoring prediction confidence scores (e.g., if the model
becomes less confident over time). Option A is performance-related; C is load; D is resource usage.



Page 3

, 10. An RTPM system uses a time-series database to store metrics. The database is configured with
a retention policy of 30 days and a downsampling rule that aggregates 1-second data points into
5-minute averages after 7 days. A data scientist needs to analyze high-frequency patterns
(sub-second) that occurred 10 days ago. Which issue will they encounter?


A. The data has been deleted due to the retention policy
B. The data has been downsampled, losing sub-second granularity
C. The data is still available at full resolution because downsampling only applies after 30 days
D. The data is stored in a separate cold storage tier with higher access latency

Answer: B
Rationale: The downsampling rule aggregates 1-second data into 5-minute averages after 7 days, so by
day 10, the sub-second resolution is lost. Option A is incorrect because retention is 30 days; C is false
because downsampling starts at 7 days; D is not mentioned.


11. In the context of real-time performance monitoring (RTPM) for a distributed microservices
architecture, which of the following best describes the primary challenge when using a push-based
metrics collection model (e.g., Prometheus Pushgateway) for ephemeral batch jobs?

A. Increased latency due to the need for service discovery and dynamic reconfiguration of scrape targets.
B. Potential for metric staleness and accumulation of stale data in the gateway if job completion fails to trigger a
cleanup.
C. Inability to aggregate metrics across multiple instances of the same job due to lack of labeling support.
D. Higher resource consumption on the monitored service because of the need to maintain a persistent
connection to the gateway.

Answer: B
Rationale: Pushgateway can retain metrics indefinitely if the pushing job does not explicitly delete them,
leading to stale data. Option A describes a pull-model challenge; C is incorrect because labeling is
supported; D is incorrect because push-based models typically have lower resource overhead on the
service.


12. A real-time predictive maintenance system uses sensor data from industrial equipment to
forecast failures. The system employs a sliding window of the last 1000 data points and updates a
logistic regression model every minute. Which of the following is the most critical design
consideration to ensure the model remains accurate over time?

A. Using a fixed random seed for model training to ensure reproducibility of results.
B. Implementing concept drift detection to trigger model retraining when the statistical properties of the data
change.
C. Increasing the sliding window size to 10,000 data points to reduce variance in the model's predictions.
D. Deploying the model on a GPU cluster to reduce training latency.

Answer: B
Rationale: Concept drift is a fundamental issue in real-time systems where data distributions change over
time, making models stale. Option A addresses reproducibility but not accuracy; C might reduce
variance but could increase latency and still not handle drift; D improves speed but does not address
accuracy if the model is outdated.




Page 4

Document information

Uploaded on
July 16, 2026
Number of pages
31
Written in
2025/2026
Type
Exam (elaborations)
Contains
Questions & answers
$28.99

Wrong document? Swap it for free Within 14 days of purchase and before downloading, you can choose a different document. You can simply spend the amount again.
Written by students who passed
Immediately available after payment
Read online or as PDF

Seller avatar
Reputation scores are based on the amount of documents a seller has sold for a fee and the reviews they have received for those documents. There are three levels: Bronze, Silver and Gold. The better the reputation, the more your can rely on the quality of the sellers work.
Studymart
4.8
(350)
Sold
105
Followers
65
Items
1376
Last sold
1 week ago



Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their tests and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can instantly pick a different document that better fits what you're looking for.

Pay as you like, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card and download your PDF document instantly.

Student with book image

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Working on your references?

Create accurate citations in APA, MLA and Harvard with our free citation generator.

Working on your references?

Frequently asked questions