Untitled Studyset

Created by JAN HAROLD FEUDO

Cloud Computing
Delivery of computing services (servers, storage, databases, networking, software) over the internet.

1/43

TermDefinition
Cloud Computing
Delivery of computing services (servers, storage, databases, networking, software) over the internet.
3 Cloud Service Models
Iaas Paas Saas
Linux - Store user acc info?
/etc/passwd
Linux - Show running processes?
ps or top
DNS
Convert domain names into IP Addresses
Protocol used for Web Traffic?
HTTP / HTTPS
Check Network Connectivity?
ping
ITIL - Incident?
Unplanned Interruption
ITIL - Goal of Incident Management?
Restore Service ASAP
Incident vs Problem?
Incident = Service Interruption Problem = Root cause of Incidents
DevOps? Main Goal?
Culture combining Development and Operations to deliver software faster and more reliably Create Scalable and Reliable Systems
CI/CD?
Continuous Integration Continuous Deployment/Delivery
SRE?
Apply SE principles to Operations and Reliability
Who Created SRE?
Google
Main Goal of SRE?
Create Scalable and Reliable Systems
SRE Principles - TOIL?
Manual, Repetitive Operational Work
Blameless Postmortems
Review Incidents without blaming individuals
Purpose of Blameless Culture?
Learn and Improve Processes
SLI
Service Layer Indicator - Measurement
SLO
Service Layer Objectives - Target Value for Service Reliability
SLA
Service Layer Agreement - Formal Agreement with Customers
Error Budget?
Amount of Acceptable failure allowed before reliability work is prioritized
Error budget Formula?
100 - SLO
When Error budget is Exhausted?
New feature releases may stop until reliability improves
3 Pillars of Observability
Metrics, Logs, Traces
Monitoring vs Observability
Monitoring: Detects issues Observability: Explains why issues happened
Which Metric Shows server CPU Usage?
CPU Utilization
Good Alert Characteristics?
Actionable Timely Relevant
Alter Fatigue?
Too many alters causing engrs to ignore them
Reliability Metrics - Avaliability Formula?
Uptime / Total Time x 100
Load Balancer?
Distribute traffics across multiple servers
On-Call
Respond to Production Incidents
What should an on-call engineer do first?
Assess impacts and restore service
Emergency Response - During Major Outage, Priority is?
Restore Service First, Root Cause analysis comes later
Why Test Reliability
Ensure systems continue working under stress or failures
Chaos Engineering?
Intentionally introducing failures to test system reliance
Goal of Chaos Engineering?
Find weaknesses before real failures occur
Chaos Monkey?
Netflix tool that randomly terminates servers
Litmus?
Kurbernetes Chaos Engineering Experiments (create small problems inside containers)
Gremlin?
Commercial Chaos Engineering Platform
Chaos Mesh
kurbernetes-native chaos testing tool (create wide variety of failures)
AIOps?
Using AI and ML to Automate IT Operations
Benefits of AIOps?
Faster Incident Detection Root Cause Analysis Predictive Monitoring