Skip to content

About

Fast, practical Linux reference covering commands, administration, troubleshooting, networking, processes, permissions, performance, Bash, systemd, logs, security, and interview questions.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

3 Commits

Folders and files

Repository files navigation

Linux-Production-Notes

Fast, practical Linux reference covering commands, administration, troubleshooting, networking, processes, permissions, performance, Bash, systemd, logs, security, and interview questions.

🐧 Linux Quick Reference

Fast last-minute notes for Linux administration, troubleshooting, networking, performance, Bash, systemd, security, and interviews.

A practical Linux reference built for learning, revision, troubleshooting, production support, and technical interviews.

This repository focuses on one question:

"The Linux system is behaving unexpectedly. How do I find out why?"


🎯 What This Repository Covers

Linux Fundamentals
        ↓
Files & Permissions
        ↓
Processes & Threads
        ↓
Memory Management
        ↓
Storage & I/O
        ↓
Networking
        ↓
systemd & Services
        ↓
Logs & Text Processing
        ↓
Bash Scripting
        ↓
Performance Analysis
        ↓
Security
        ↓
Kernel & /proc
        ↓
Production Troubleshooting
        ↓
Interview Preparation

⚑ 60-Second Linux Cheat Sheet

System Information

uname -a
hostname
hostnamectl
uptime
date
whoami
id

CPU

lscpu
top
htop
mpstat
uptime

Memory

free -h
vmstat 1
cat /proc/meminfo

Disk

df -h
du -sh *
lsblk
blkid
mount

Processes

ps aux
ps -ef
top
pgrep <name>
pidof <name>
kill <PID>
kill -9 <PID>

Networking

ip addr
ip route
ss -tulpn
ping <host>
traceroute <host>
curl <url>
dig <domain>
nslookup <domain>

Logs

journalctl
journalctl -u <service>
journalctl -f
tail -f /var/log/<file>
grep "ERROR" <file>

Services

systemctl status <service>
systemctl start <service>
systemctl stop <service>
systemctl restart <service>
systemctl enable <service>
systemctl disable <service>

Open Files

lsof
lsof -i :8080
lsof -p <PID>

🧠 The Linux Troubleshooting Mindset

Do not randomly execute commands until something looks wrong.

Start with:

What is the symptom?
        ↓
What changed?
        ↓
Is the issue reproducible?
        ↓
Is it CPU?
Memory?
Disk?
Network?
Process?
Service?
Application?
        ↓
Collect evidence
        ↓
Form hypothesis
        ↓
Test hypothesis
        ↓
Fix
        ↓
Verify
        ↓
Prevent recurrence

For performance investigations, the USE Method is a useful framework:

Utilization β†’ Saturation β†’ Errors

Check resources systematically rather than jumping between unrelated commands. Brendan Gregg documents this methodology specifically for identifying resource bottlenecks and errors.


πŸ“š Core Topics

01 β€” Linux Basics

  • Linux architecture
  • Kernel vs user space
  • Shell
  • Filesystem hierarchy
  • Absolute vs relative paths
  • Environment variables
  • stdin / stdout / stderr
  • Pipes
  • Redirection
  • Wildcards
  • Command substitution

02 β€” Files & Permissions

Understand:

r = read
w = write
x = execute

Example:

-rwxr-xr--

Breakdown:

Owner   Group   Others
rwx     r-x     r--

Important commands:

chmod
chown
chgrp
umask
getfacl
setfacl

03 β€” Processes

A process is a running instance of a program.

Useful commands:

ps
top
htop
pgrep
pkill
kill
killall
nice
renice
jobs
bg
fg

Important concepts:

PID
PPID
Process states
Foreground/background
Signals
Zombie
Orphan
Daemon

04 β€” Signals

Common signals:

SIGTERM  β†’ request graceful termination
SIGKILL  β†’ force termination
SIGINT   β†’ interrupt
SIGHUP   β†’ hangup
SIGSTOP  β†’ stop
SIGCONT  β†’ continue

Example:

kill -TERM <PID>

Use SIGKILL carefully:

kill -9 <PID>

It does not give the process an opportunity to clean up normally.


05 β€” Memory

Important concepts:

Physical memory
Virtual memory
Pages
Swap
Cache
Buffers
OOM Killer
Page faults

Useful commands:

free -h
vmstat 1
top
ps aux --sort=-%mem
cat /proc/meminfo

Quick diagnosis

High memory usage
      ↓
free -h
      ↓
Is available memory low?
      ↓
Check swap
      ↓
vmstat
      ↓
Find memory-heavy processes
      ↓
Investigate application

06 β€” Storage

Useful commands:

df -h
du -sh *
lsblk
mount
findmnt
blkid
lsof

Important distinction

df
β†’ filesystem space

du
β†’ space consumed by files/directories

If:

df says disk is full
du doesn't explain it

investigate deleted-but-open files:

lsof +L1

07 β€” Networking

Core concepts:

IP
MAC
ARP
DNS
TCP
UDP
Ports
Routing
Sockets
NAT
HTTP
HTTPS
TLS

Useful commands:

ip addr
ip route
ss -tulpn
ping
traceroute
curl
dig
nslookup

Port troubleshooting

ss -lntp

Check whether a service is listening.

Then:

curl localhost:<PORT>

Then investigate:

Application
    ↓
Process
    ↓
Socket
    ↓
Firewall
    ↓
Route
    ↓
Remote server

08 β€” DNS Troubleshooting

When a hostname fails:

dig example.com

Check:

DNS resolution
        ↓
IP address
        ↓
Routing
        ↓
TCP connectivity
        ↓
Application protocol

Do not assume a "website is down" when DNS is the actual problem.


09 β€” systemd

Important commands:

systemctl status nginx
systemctl start nginx
systemctl stop nginx
systemctl restart nginx
systemctl enable nginx
systemctl disable nginx

Logs:

journalctl -u nginx
journalctl -u nginx -f
journalctl -b
journalctl -p err

Typical service investigation:

Service failed
      ↓
systemctl status
      ↓
journalctl
      ↓
Check configuration
      ↓
Check dependencies
      ↓
Check port
      ↓
Check permissions
      ↓
Restart
      ↓
Verify

10 β€” Logs

Important tools:

cat
less
tail
head
grep
awk
sed
cut
sort
uniq
wc
xargs

Real-world example:

grep "ERROR" application.log | tail -50

Count error types:

grep "ERROR" application.log \
  | awk '{print $NF}' \
  | sort \
  | uniq -c \
  | sort -nr

11 β€” Bash

Core topics:

Variables
Arguments
Exit codes
Conditions
Loops
Functions
Arrays
Command substitution
Pipes
Redirection
Signals
Trap
Cron

Example:

#!/bin/bash

if systemctl is-active --quiet nginx; then
    echo "nginx is running"
else
    echo "nginx is NOT running"
fi

Always check exit codes when writing operational scripts:

echo $?

12 β€” Performance Troubleshooting

Start with a structured investigation.

CPU

top
mpstat -P ALL 1
pidstat -u 1

Memory

free -h
vmstat 1
pidstat -r 1

Disk I/O

iostat -xz 1
pidstat -d 1

Network

ss -s
sar -n DEV 1
ip -s link

System calls

strace -p <PID>

Advanced profiling

perf

For deeper performance work, investigate:

CPU profiling
Off-CPU analysis
Flame graphs
perf events
eBPF
ftrace

Brendan Gregg's Linux performance material provides deeper references for these techniques.


13 β€” /proc

/proc is a pseudo-filesystem exposing kernel and process information to user space.

Examples:

cat /proc/cpuinfo
cat /proc/meminfo
cat /proc/loadavg
cat /proc/uptime

For a process:

cat /proc/<PID>/status
ls -l /proc/<PID>/fd

Understanding /proc makes many Linux tools easier to reason about rather than treating them as magic commands.


14 β€” Security

Important topics:

Users
Groups
sudo
SSH
File permissions
ACL
SSH keys
Firewall
Processes
Capabilities
SELinux/AppArmor

Useful commands:

id
who
w
last
sudo
ssh
ss

Never blindly execute commands copied from the internet as root.


🚨 Production Troubleshooting Scenarios

Scenario 1 β€” CPU is 100%

Start:

uptime
top
ps aux --sort=-%cpu | head
mpstat -P ALL 1

Questions:

  • Is one process responsible?
  • Is CPU evenly distributed?
  • Is there high load but low CPU utilization?
  • Is the process actually CPU-bound?
  • Did a deployment happen recently?

Scenario 2 β€” Server is slow

Don't immediately blame CPU.

Check:

CPU
Memory
Disk I/O
Network
Load
Processes
Application
External dependencies

Use the USE framework where appropriate.


Scenario 3 β€” Disk is full

df -h
du -xhd1 /

Then:

find /var -type f -size +1G

Also investigate:

lsof +L1

Scenario 4 β€” Port is not reachable

ss -lntp

Check:

Is process running?
       ↓
Is application listening?
       ↓
Is it bound to correct IP?
       ↓
Firewall?
       ↓
Routing?
       ↓
Remote connectivity?

Scenario 5 β€” Service won't start

systemctl status <service>
journalctl -u <service> -n 100

Then check:

Configuration
Permissions
Dependencies
Port conflicts
Environment variables
Disk space
Resource limits

🎯 Most Asked Linux Interview Questions

Fundamentals

  1. What is Linux?
  2. What is the Linux kernel?
  3. Kernel vs shell?
  4. Linux vs Unix?
  5. What happens when you execute a command?
  6. What is a process?
  7. What is a thread?
  8. What is a daemon?
  9. What is a system call?
  10. What is /proc?

Processes

  1. Process vs thread?
  2. What is a PID?
  3. What is PPID?
  4. What is a zombie process?
  5. What is an orphan process?
  6. What happens when you run kill -9?
  7. SIGTERM vs SIGKILL?
  8. How do you find a process?
  9. How do you kill a process?
  10. How do you find which process is consuming CPU?

Memory

  1. What is virtual memory?
  2. What is swap?
  3. What is paging?
  4. What is a page fault?
  5. What is OOM Killer?
  6. How do you check memory usage?
  7. Why can free memory appear low even when the system is healthy?
  8. How would you troubleshoot high memory usage?

Filesystem

  1. df vs du?
  2. Hard link vs symbolic link?
  3. What is an inode?
  4. What happens when a file is deleted?
  5. How can disk space remain consumed after deleting a file?
  6. What is /etc?
  7. What is /var?
  8. What is /tmp?
  9. What is /proc?
  10. What is /dev?

Permissions

  1. Explain 755.
  2. Explain 644.
  3. What does chmod do?
  4. chmod vs chown?
  5. What is umask?
  6. What is ACL?
  7. How do you troubleshoot "Permission denied"?

Networking

  1. TCP vs UDP?
  2. What is a port?
  3. What is a socket?
  4. What does ss show?
  5. ping vs traceroute?
  6. What is DNS?
  7. How do you troubleshoot DNS?
  8. How do you check listening ports?
  9. How do you find which process owns port 8080?
  10. How would you troubleshoot a server that is unreachable?

systemd

  1. How do you start a service?
  2. How do you stop a service?
  3. How do you restart a service?
  4. How do you enable a service at boot?
  5. How do you view service logs?
  6. What is journalctl?
  7. How do you troubleshoot a failed service?

πŸ”₯ Scenario-Based Interview Questions

Q1. Server CPU suddenly reaches 100%. What do you do?

Don't answer:

"I will run top."

A stronger answer:

1. Confirm the symptom.
2. Check load average.
3. Identify CPU-consuming processes.
4. Determine whether CPU usage is user/system/iowait.
5. Check per-core utilization.
6. Inspect the responsible process.
7. Check recent deployments/configuration changes.
8. Determine root cause.
9. Mitigate safely.
10. Verify recovery.
11. Document prevention.

Q2. Server has 90% memory usage. Is that automatically a problem?

No.

Investigate:

free -h
vmstat 1
ps aux --sort=-%mem

Look at:

available memory
swap activity
reclaimable cache
OOM events
application behavior

High memory utilization by itself does not establish that the system is unhealthy.


Q3. Application is down but process is running.

Investigate:

Process
 ↓
Listening socket
 ↓
Application logs
 ↓
Local request
 ↓
Firewall
 ↓
Network path
 ↓
Dependency

Useful commands:

ps
ss
curl
journalctl
grep
lsof

Q4. Disk says 100% full but du doesn't explain it.

Investigate:

lsof +L1

A process may still have an open file descriptor for a deleted file.


🧰 Essential Commands by Problem

Problem First Commands
CPU top, ps, mpstat, pidstat
Memory free, vmstat, ps
Disk space df, du
Disk I/O iostat, pidstat
Process ps, pgrep, top
Port ss, lsof
Network ip, ss, ping, curl
DNS dig, nslookup
Service systemctl, journalctl
Logs grep, tail, less
Kernel dmesg, /proc
Syscalls strace
Performance perf, vmstat, iostat

🧭 Troubleshooting Decision Tree

             INCIDENT
                 β”‚
                 β–Ό
        What is the symptom?
                 β”‚
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β–Ό         β–Ό         β–Ό
      CPU      Memory     Network
       β”‚         β”‚         β”‚
      top      free       ip
      ps       vmstat     ss
      mpstat   ps         ping
       β”‚         β”‚         β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                 β–Ό
              Storage
                 β”‚
             df / du
             iostat
                 β”‚
                 β–Ό
              Process
                 β”‚
             ps / lsof
                 β”‚
                 β–Ό
              Service
                 β”‚
       systemctl / journalctl
                 β”‚
                 β–Ό
              Logs
                 β”‚
        grep / awk / sed
                 β”‚
                 β–Ό
            Root Cause
                 β”‚
                 β–Ό
              Fix
                 β”‚
                 β–Ό
             Verify
                 β”‚
                 β–Ό
            Prevent

πŸ“– Recommended References

Official / Primary Documentation

Performance & Troubleshooting

The /proc filesystem exposes kernel data structures and process/system information, making it an important source for understanding how Linux exposes runtime state.


πŸ“š Recommended Book

Systems Performance: Enterprise and the Cloud β€” Brendan Gregg

Use it as a deeper reference for:

CPU
Memory
Disks
Networking
Filesystems
Kernel
Performance methodology
Observability

The Linux USE checklist is included as an appendix in the second edition.


πŸ§ͺ Hands-On Labs

This repository will eventually include practical labs for:

  • High CPU
  • Memory pressure
  • Disk full
  • Disk I/O bottleneck
  • Zombie processes
  • Hanging processes
  • Port conflicts
  • DNS failure
  • Network latency
  • Service failure
  • Permission problems
  • Log analysis
  • SSH problems
  • File descriptor exhaustion
  • OOM conditions

πŸš€ Goal

The goal is not to memorize 200 Linux commands.

The goal is to be able to look at a Linux system and reason:

What is happening?
       ↓
Where is it happening?
       ↓
Why is it happening?
       ↓
How can I prove it?
       ↓
How do I fix it?
       ↓
How do I prevent it?

Learn the command. Understand the system. Diagnose the problem.


πŸ‘¨β€πŸ’» Author

Prateek

Software Engineering β€’ Linux β€’ Java β€’ SQL β€’ Cloud β€’ Systems

Learning by building, troubleshooting, and documenting.

About

Fast, practical Linux reference covering commands, administration, troubleshooting, networking, processes, permissions, performance, Bash, systemd, logs, security, and interview questions.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors