Untitled Studyset
Created by JAN HAROLD FEUDO
Cloud Computing
Delivery of computing services (servers, storage, databases, networking, software) over the internet.
| Term | Definition |
|---|---|
Cloud Computing | Delivery of computing services (servers, storage, databases, networking, software) over the internet. |
3 Cloud Service Models | Iaas
Paas
Saas |
Linux - Store user acc info? | /etc/passwd |
Linux - Show running processes? | ps or top |
DNS | Convert domain names into IP Addresses |
Protocol used for Web Traffic? | HTTP / HTTPS |
Check Network Connectivity? | ping |
ITIL - Incident? | Unplanned Interruption |
ITIL - Goal of Incident Management? | Restore Service ASAP |
Incident vs Problem? | Incident = Service Interruption
Problem = Root cause of Incidents
|
DevOps? Main Goal? | Culture combining Development and Operations to deliver software faster and more reliably
Create Scalable and Reliable Systems |
CI/CD? | Continuous Integration
Continuous Deployment/Delivery |
SRE? | Apply SE principles to Operations and Reliability
|
Who Created SRE? | Google |
Main Goal of SRE? | Create Scalable and Reliable Systems |
SRE Principles - TOIL? | Manual, Repetitive Operational Work |
Blameless Postmortems | Review Incidents without blaming individuals |
Purpose of Blameless Culture? | Learn and Improve Processes |
SLI | Service Layer Indicator - Measurement |
SLO | Service Layer Objectives - Target Value for Service Reliability |
SLA | Service Layer Agreement - Formal Agreement with Customers |
Error Budget? | Amount of Acceptable failure allowed before reliability work is prioritized |
Error budget Formula? | 100 - SLO |
When Error budget is Exhausted? | New feature releases may stop until reliability improves
|
3 Pillars of Observability | Metrics, Logs, Traces |
Monitoring vs Observability | Monitoring: Detects issues
Observability: Explains why issues happened |
Which Metric Shows server CPU Usage? | CPU Utilization
|
Good Alert Characteristics? | Actionable
Timely
Relevant |
Alter Fatigue? | Too many alters causing engrs to ignore them |
Reliability Metrics - Avaliability Formula? | Uptime / Total Time x 100 |
Load Balancer? | Distribute traffics across multiple servers |
On-Call | Respond to Production Incidents |
What should an on-call engineer do first? | Assess impacts and restore service |
Emergency Response - During Major Outage, Priority is? | Restore Service First, Root Cause analysis comes later |
Why Test Reliability | Ensure systems continue working under stress or failures |
Chaos Engineering? | Intentionally introducing failures to test system reliance |
Goal of Chaos Engineering? | Find weaknesses before real failures occur |
Chaos Monkey? | Netflix tool that randomly terminates servers |
Litmus? | Kurbernetes Chaos Engineering Experiments (create small problems inside containers) |
Gremlin? | Commercial Chaos Engineering Platform |
Chaos Mesh | kurbernetes-native chaos testing tool (create wide variety of failures)
|
AIOps? | Using AI and ML to Automate IT Operations |
Benefits of AIOps? | Faster Incident Detection
Root Cause Analysis
Predictive Monitoring |