Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Operations and monitoring in DevSecOps determine whether teams can detect, withstand, and recover from threats in production systems and tenant environments. Strong operational practices help teams prove resilience under real-world conditions.
This capability area describes how an organization protects its tenants and production systems, monitors for and detects threats, and responds to and remediates incidents. It's where resilience is tested continuously in live environments.
As organizations mature, operations and monitoring move from minimal, manual practices toward fortified operations supported by prediction. Production systems become more hardened, access is controlled more tightly, monitoring evolves from simple metrics to automated anomaly detection and forecasting, and incident response matures from unplanned reaction to AI-assisted execution and regular drills.
Core capability areas
Operations and monitoring in DevSecOps focuses on three sub-capabilities:
Protecting tenants and production systems: How production systems, devices, identities, and images are hardened and controlled.
Monitoring and detection of threats: How logging, metrics, and detection surface security and reliability issues.
Incident response and remediation: How the organization plans for, drills, and responds to incidents.
Stages
Operations and monitoring progresses through five stages of maturity. These stages show how an organization applies DevSecOps principles in day-to-day operations, from manual support of service continuity to fortified operations that can anticipate issues before they happen.
| Overall stage | Stage name | What it looks like |
|---|---|---|
| Ad-hoc | Minimal | Manual or outdated tools make it difficult to support service continuity and stability. |
| Initiating | Instantiated | Establishing operations improves performance, although capabilities remain limited and mostly default. |
| Orchestrating | Robust | Centering operations on security and reliability expands insight and enables earlier issue identification. |
| Streamlining | Resilient | Automation and advanced data synthesis strengthen operations capabilities. |
| Pioneering | Fortified | Predictive capabilities help teams anticipate issues before they happen. |
Minimal (Ad-hoc)
When teams rely on manual or outdated tools and services, it becomes difficult to support service continuity and stability effectively.
Protecting tenants and production systems: Unused, aging, and legacy systems remain in place and vulnerable. There are no requirements for device health when accessing systems.
Monitoring and detection of threats: Teams rely on centralized system logging, simple application metrics, simple budget metrics, and simple system metrics.
Incident response and remediation: A lack of planned response capabilities delays patching and release updates.
Instantiated (Initiating)
Establishing product and service operations improves performance, but often through default options with limited capability.
Protecting tenants and production systems: Unused, aging, and legacy systems are removed. Access to systems is restricted to secured, managed, and healthy devices.
Monitoring and detection of threats: Potential security-related events are logged, with alerting based on manually defined suspicious patterns. Cost monitoring and metrics and log visualization are in place.
Incident response and remediation: A basic response strategy is established and supports initial incident handling.
Robust (Orchestrating)
Centering product and service operations on security and reliability requires broader insight gathering. Earlier issue identification leads to more reliable services.
Protecting tenants and production systems: Patch policy is defined, automated pull requests are used for patches, and lateral identity movement across tenants is prevented.
Monitoring and detection of threats: Data is centralized and unified with advanced availability and stability metrics, auditing of system events, deactivation of unused metrics, grouping of metrics, and targeted alerting.
Incident response and remediation: Simple business continuity and disaster recovery (BCDR) practices are in place for critical components, and incident analysis becomes a routine activity.
Resilient (Streamlining)
Investing in automation and advanced data collection and synthesis helps organizations strengthen product and service operations.
Protecting tenants and production systems: Automated pull requests are merged automatically, and base images are built nightly. A maximum lifetime for images helps reduce incident impact. Continuous least-privilege access is enforced, and intrusion detection systems (IDS) are used.
Monitoring and detection of threats: Advanced application metrics, coverage and control metrics, and defense metrics are collected. Automated anomaly detection, analysis, and alerting are used, including coverage for model inversion and adversarial attacks.
Incident response and remediation: Regular incident drills and advanced incident response protocols are in place for all systems. Automated isolation and mitigation controls are also in place.
Fortified (Pioneering)
Building predictive capabilities into strong product and service operations allows organizations to anticipate issues before they happen.
Protecting tenants and production systems: A short maximum lifetime for images further reduces incident impact.
Monitoring and detection of threats: Advanced pattern detection and forecasting are used to discover and mitigate risks before they can be exploited.
Incident response and remediation: AI-driven assistance helps target threats, assess activity, and provide threat intelligence in real time.