Fast, practical Linux reference covering commands, administration, troubleshooting, networking, processes, permissions, performance, Bash, systemd, logs, security, and interview questions.
Fast last-minute notes for Linux administration, troubleshooting, networking, performance, Bash, systemd, security, and interviews.
A practical Linux reference built for learning, revision, troubleshooting, production support, and technical interviews.
This repository focuses on one question:
"The Linux system is behaving unexpectedly. How do I find out why?"
Linux Fundamentals
β
Files & Permissions
β
Processes & Threads
β
Memory Management
β
Storage & I/O
β
Networking
β
systemd & Services
β
Logs & Text Processing
β
Bash Scripting
β
Performance Analysis
β
Security
β
Kernel & /proc
β
Production Troubleshooting
β
Interview Preparation
uname -a
hostname
hostnamectl
uptime
date
whoami
idlscpu
top
htop
mpstat
uptimefree -h
vmstat 1
cat /proc/meminfodf -h
du -sh *
lsblk
blkid
mountps aux
ps -ef
top
pgrep <name>
pidof <name>
kill <PID>
kill -9 <PID>ip addr
ip route
ss -tulpn
ping <host>
traceroute <host>
curl <url>
dig <domain>
nslookup <domain>journalctl
journalctl -u <service>
journalctl -f
tail -f /var/log/<file>
grep "ERROR" <file>systemctl status <service>
systemctl start <service>
systemctl stop <service>
systemctl restart <service>
systemctl enable <service>
systemctl disable <service>lsof
lsof -i :8080
lsof -p <PID>Do not randomly execute commands until something looks wrong.
Start with:
What is the symptom?
β
What changed?
β
Is the issue reproducible?
β
Is it CPU?
Memory?
Disk?
Network?
Process?
Service?
Application?
β
Collect evidence
β
Form hypothesis
β
Test hypothesis
β
Fix
β
Verify
β
Prevent recurrence
For performance investigations, the USE Method is a useful framework:
Utilization β Saturation β Errors
Check resources systematically rather than jumping between unrelated commands. Brendan Gregg documents this methodology specifically for identifying resource bottlenecks and errors.
- Linux architecture
- Kernel vs user space
- Shell
- Filesystem hierarchy
- Absolute vs relative paths
- Environment variables
- stdin / stdout / stderr
- Pipes
- Redirection
- Wildcards
- Command substitution
Understand:
r = read
w = write
x = execute
Example:
-rwxr-xr--Breakdown:
Owner Group Others
rwx r-x r--
Important commands:
chmod
chown
chgrp
umask
getfacl
setfaclA process is a running instance of a program.
Useful commands:
ps
top
htop
pgrep
pkill
kill
killall
nice
renice
jobs
bg
fgImportant concepts:
PID
PPID
Process states
Foreground/background
Signals
Zombie
Orphan
Daemon
Common signals:
SIGTERM β request graceful termination
SIGKILL β force termination
SIGINT β interrupt
SIGHUP β hangup
SIGSTOP β stop
SIGCONT β continue
Example:
kill -TERM <PID>Use SIGKILL carefully:
kill -9 <PID>It does not give the process an opportunity to clean up normally.
Important concepts:
Physical memory
Virtual memory
Pages
Swap
Cache
Buffers
OOM Killer
Page faults
Useful commands:
free -h
vmstat 1
top
ps aux --sort=-%mem
cat /proc/meminfoHigh memory usage
β
free -h
β
Is available memory low?
β
Check swap
β
vmstat
β
Find memory-heavy processes
β
Investigate application
Useful commands:
df -h
du -sh *
lsblk
mount
findmnt
blkid
lsofdf
β filesystem space
du
β space consumed by files/directories
If:
df says disk is full
du doesn't explain it
investigate deleted-but-open files:
lsof +L1Core concepts:
IP
MAC
ARP
DNS
TCP
UDP
Ports
Routing
Sockets
NAT
HTTP
HTTPS
TLS
Useful commands:
ip addr
ip route
ss -tulpn
ping
traceroute
curl
dig
nslookupss -lntpCheck whether a service is listening.
Then:
curl localhost:<PORT>Then investigate:
Application
β
Process
β
Socket
β
Firewall
β
Route
β
Remote server
When a hostname fails:
dig example.comCheck:
DNS resolution
β
IP address
β
Routing
β
TCP connectivity
β
Application protocol
Do not assume a "website is down" when DNS is the actual problem.
Important commands:
systemctl status nginx
systemctl start nginx
systemctl stop nginx
systemctl restart nginx
systemctl enable nginx
systemctl disable nginxLogs:
journalctl -u nginx
journalctl -u nginx -f
journalctl -b
journalctl -p errTypical service investigation:
Service failed
β
systemctl status
β
journalctl
β
Check configuration
β
Check dependencies
β
Check port
β
Check permissions
β
Restart
β
Verify
Important tools:
cat
less
tail
head
grep
awk
sed
cut
sort
uniq
wc
xargsReal-world example:
grep "ERROR" application.log | tail -50Count error types:
grep "ERROR" application.log \
| awk '{print $NF}' \
| sort \
| uniq -c \
| sort -nrCore topics:
Variables
Arguments
Exit codes
Conditions
Loops
Functions
Arrays
Command substitution
Pipes
Redirection
Signals
Trap
Cron
Example:
#!/bin/bash
if systemctl is-active --quiet nginx; then
echo "nginx is running"
else
echo "nginx is NOT running"
fiAlways check exit codes when writing operational scripts:
echo $?Start with a structured investigation.
top
mpstat -P ALL 1
pidstat -u 1free -h
vmstat 1
pidstat -r 1iostat -xz 1
pidstat -d 1ss -s
sar -n DEV 1
ip -s linkstrace -p <PID>perfFor deeper performance work, investigate:
CPU profiling
Off-CPU analysis
Flame graphs
perf events
eBPF
ftrace
Brendan Gregg's Linux performance material provides deeper references for these techniques.
/proc is a pseudo-filesystem exposing kernel and process information to user space.
Examples:
cat /proc/cpuinfo
cat /proc/meminfo
cat /proc/loadavg
cat /proc/uptimeFor a process:
cat /proc/<PID>/status
ls -l /proc/<PID>/fdUnderstanding /proc makes many Linux tools easier to reason about rather than treating them as magic commands.
Important topics:
Users
Groups
sudo
SSH
File permissions
ACL
SSH keys
Firewall
Processes
Capabilities
SELinux/AppArmor
Useful commands:
id
who
w
last
sudo
ssh
ssNever blindly execute commands copied from the internet as root.
Start:
uptime
top
ps aux --sort=-%cpu | head
mpstat -P ALL 1Questions:
- Is one process responsible?
- Is CPU evenly distributed?
- Is there high load but low CPU utilization?
- Is the process actually CPU-bound?
- Did a deployment happen recently?
Don't immediately blame CPU.
Check:
CPU
Memory
Disk I/O
Network
Load
Processes
Application
External dependencies
Use the USE framework where appropriate.
df -h
du -xhd1 /Then:
find /var -type f -size +1GAlso investigate:
lsof +L1ss -lntpCheck:
Is process running?
β
Is application listening?
β
Is it bound to correct IP?
β
Firewall?
β
Routing?
β
Remote connectivity?
systemctl status <service>
journalctl -u <service> -n 100Then check:
Configuration
Permissions
Dependencies
Port conflicts
Environment variables
Disk space
Resource limits
- What is Linux?
- What is the Linux kernel?
- Kernel vs shell?
- Linux vs Unix?
- What happens when you execute a command?
- What is a process?
- What is a thread?
- What is a daemon?
- What is a system call?
- What is
/proc?
- Process vs thread?
- What is a PID?
- What is PPID?
- What is a zombie process?
- What is an orphan process?
- What happens when you run
kill -9? - SIGTERM vs SIGKILL?
- How do you find a process?
- How do you kill a process?
- How do you find which process is consuming CPU?
- What is virtual memory?
- What is swap?
- What is paging?
- What is a page fault?
- What is OOM Killer?
- How do you check memory usage?
- Why can
freememory appear low even when the system is healthy? - How would you troubleshoot high memory usage?
dfvsdu?- Hard link vs symbolic link?
- What is an inode?
- What happens when a file is deleted?
- How can disk space remain consumed after deleting a file?
- What is
/etc? - What is
/var? - What is
/tmp? - What is
/proc? - What is
/dev?
- Explain
755. - Explain
644. - What does
chmoddo? chmodvschown?- What is
umask? - What is ACL?
- How do you troubleshoot "Permission denied"?
- TCP vs UDP?
- What is a port?
- What is a socket?
- What does
ssshow? pingvstraceroute?- What is DNS?
- How do you troubleshoot DNS?
- How do you check listening ports?
- How do you find which process owns port 8080?
- How would you troubleshoot a server that is unreachable?
- How do you start a service?
- How do you stop a service?
- How do you restart a service?
- How do you enable a service at boot?
- How do you view service logs?
- What is
journalctl? - How do you troubleshoot a failed service?
Don't answer:
"I will run
top."
A stronger answer:
1. Confirm the symptom.
2. Check load average.
3. Identify CPU-consuming processes.
4. Determine whether CPU usage is user/system/iowait.
5. Check per-core utilization.
6. Inspect the responsible process.
7. Check recent deployments/configuration changes.
8. Determine root cause.
9. Mitigate safely.
10. Verify recovery.
11. Document prevention.
No.
Investigate:
free -h
vmstat 1
ps aux --sort=-%memLook at:
available memory
swap activity
reclaimable cache
OOM events
application behavior
High memory utilization by itself does not establish that the system is unhealthy.
Investigate:
Process
β
Listening socket
β
Application logs
β
Local request
β
Firewall
β
Network path
β
Dependency
Useful commands:
ps
ss
curl
journalctl
grep
lsofInvestigate:
lsof +L1A process may still have an open file descriptor for a deleted file.
| Problem | First Commands |
|---|---|
| CPU | top, ps, mpstat, pidstat |
| Memory | free, vmstat, ps |
| Disk space | df, du |
| Disk I/O | iostat, pidstat |
| Process | ps, pgrep, top |
| Port | ss, lsof |
| Network | ip, ss, ping, curl |
| DNS | dig, nslookup |
| Service | systemctl, journalctl |
| Logs | grep, tail, less |
| Kernel | dmesg, /proc |
| Syscalls | strace |
| Performance | perf, vmstat, iostat |
INCIDENT
β
βΌ
What is the symptom?
β
βββββββββββΌββββββββββ
βΌ βΌ βΌ
CPU Memory Network
β β β
top free ip
ps vmstat ss
mpstat ps ping
β β β
βββββββββββΌββββββββββ
βΌ
Storage
β
df / du
iostat
β
βΌ
Process
β
ps / lsof
β
βΌ
Service
β
systemctl / journalctl
β
βΌ
Logs
β
grep / awk / sed
β
βΌ
Root Cause
β
βΌ
Fix
β
βΌ
Verify
β
βΌ
Prevent
- Linux Kernel Documentation
- Linux man-pages
- GNU Coreutils
- Bash Reference Manual
- systemd Documentation
- OpenSSH Documentation
- Brendan Gregg β Linux Performance
- Brendan Gregg β Performance Methodology
- Brendan Gregg β USE Method
- Brendan Gregg β perf examples
- Linux
proc(5)documentation
The /proc filesystem exposes kernel data structures and process/system information, making it an important source for understanding how Linux exposes runtime state.
Systems Performance: Enterprise and the Cloud β Brendan Gregg
Use it as a deeper reference for:
CPU
Memory
Disks
Networking
Filesystems
Kernel
Performance methodology
Observability
The Linux USE checklist is included as an appendix in the second edition.
This repository will eventually include practical labs for:
- High CPU
- Memory pressure
- Disk full
- Disk I/O bottleneck
- Zombie processes
- Hanging processes
- Port conflicts
- DNS failure
- Network latency
- Service failure
- Permission problems
- Log analysis
- SSH problems
- File descriptor exhaustion
- OOM conditions
The goal is not to memorize 200 Linux commands.
The goal is to be able to look at a Linux system and reason:
What is happening?
β
Where is it happening?
β
Why is it happening?
β
How can I prove it?
β
How do I fix it?
β
How do I prevent it?
Learn the command. Understand the system. Diagnose the problem.
Prateek
Software Engineering β’ Linux β’ Java β’ SQL β’ Cloud β’ Systems
Learning by building, troubleshooting, and documenting.