diff --git a/questions/README.md b/questions/README.md index 54ab95a..44fb241 100644 --- a/questions/README.md +++ b/questions/README.md @@ -2,7 +2,7 @@ ## Table of Contents -- [1- Python uses a Global Interpreter Lock. Does that mean it doesn’t use actual threads?](#1--python-uses-a-global-interpreter-lock-does-that-mean-it-doesnt-use-actual-threads) +- [1- Python uses a Global Interpreter Lock. Does that mean it doesn't use actual threads?](#1--python-uses-a-global-interpreter-lock-does-that-mean-it-doesnt-use-actual-threads) - [2- Is it possible to have a producer thread reading from the network and a consumer thread writing to a file work in parallel? What about the GIL?](#2--is-it-possible-to-have-a-producer-thread-reading-from-the-network-and-a-consumer-thread-writing-to-a-file-work-in-parallel-what-about-the-gil) - [3- What will be the output of the following code in each step?](#3--what-will-be-the-output-of-the-following-code-in-each-step) - [4- Why are functions considered first-class objects in Python?](#4--why-are-functions-considered-first-class-objects-in-python) @@ -16,23 +16,23 @@ - [12- What will be the output of the following code?](#12--what-will-be-the-output-of-the-following-code) - [13- A palindromic number reads the same both ways. The largest palindrome made from the product of two 2-digit numbers is 9009 = 91 × 99. Find the largest palindrome made from the product of two 3-digit numbers](#13--a-palindromic-number-reads-the-same-both-ways-the-largest-palindrome-made-from-the-product-of-two-2-digit-numbers-is-9009--91--99-find-the-largest-palindrome-made-from-the-product-of-two-3-digit-numbers) - [14- What is skeleton code in Python?](#14--what-is-skeleton-code-in-python) -- [15- In Python classes, what is the difference between class methods and static methods? and when to use them](#15--in-python-classes-what-is-the-difference-between-class-methods-and-static-methods-and-when-to-use-them) +- [15- In Python classes, what is the difference between class methods and static methods, and when should you use each?](#15--in-python-classes-what-is-the-difference-between-class-methods-and-static-methods-and-when-should-you-use-each) - [16- Please explain the following results of the code executed on a Python shell interpreter](#16--please-explain-the-following-results-of-the-code-executed-on-a-python-shell-interpreter) -- [17- In object-oriented programming, there is a concept called abstract classes. How to implement it?](#17--in-object-oriented-programming-there-is-a-concept-called-abstract-classes-how-to-implement-it) -- [18- What are `*args` and `**kwargs` in Python](#18--what-are-args-and-kwargs-in-python) +- [17- In object-oriented programming, there is a concept called abstract classes. How do you implement one in Python?](#17--in-object-oriented-programming-there-is-a-concept-called-abstract-classes-how-do-you-implement-one-in-python) +- [18- What are `*args` and `**kwargs` in Python?](#18--what-are-args-and-kwargs-in-python) - [19- What is the difference between tuples, sets, and lists in Python?](#19--what-is-the-difference-between-tuples-sets-and-lists-in-python) - [20- What are pickling and unpickling in Python?](#20--what-are-pickling-and-unpickling-in-python) - [21- Does Python support multiple inheritance?](#21--does-python-support-multiple-inheritance) -- [22- What are the pitfalls and problems of Python language?](#22--what-are-the-pitfalls-and-problems-of-python-language) -- [23- How to achieve multithreading in Python?](#23--how-to-achieve-multithreading-in-python) +- [22- What are the pitfalls and problems of the Python language?](#22--what-are-the-pitfalls-and-problems-of-the-python-language) +- [23- How do you achieve multithreading in Python?](#23--how-do-you-achieve-multithreading-in-python) - [24- What is the use of `with` in Python?](#24--what-is-the-use-of-with-in-python) - [25- How are `.py`, `.pyi`, `.pyd`, and `.pyc` files different?](#25--how-are-py-pyi-pyd-and-pyc-files-different) - [26- What are decorators in Python?](#26--what-are-decorators-in-python) -- [27- How to use `self` in Python?](#27--how-to-use-self-in-python) +- [27- How do you use `self` in Python?](#27--how-do-you-use-self-in-python) - [28- What are namespaces in Python?](#28--what-are-namespaces-in-python) -- [29- What is PEP?](#29--what-is-pep) +- [29- What is a PEP?](#29--what-is-a-pep) - [30- What are dunder methods in Python?](#30--what-are-dunder-methods-in-python) -- [31- What does `super` do in Python? and what is the difference between `super().__init__()` and explicit `superclass.__init__()`](#31--what-does-super-do-in-python-and-what-is-the-difference-between-super__init__-and-explicit-superclass__init__) +- [31- What does `super` do in Python, and what is the difference between `super().__init__()` and an explicit `superclass.__init__()` call?](#31--what-does-super-do-in-python-and-what-is-the-difference-between-super__init__-and-an-explicit-superclass__init__-call) - [32- What is a property decorator in Python?](#32--what-is-a-property-decorator-in-python) - [33- What is the difference between Cython and CPython?](#33--what-is-the-difference-between-cython-and-cpython) - [34- Specify the difference between local and global variables in Python](#34--specify-the-difference-between-local-and-global-variables-in-python) @@ -40,80 +40,80 @@ - [36- What are Python generators?](#36--what-are-python-generators) - [37- What is the difference between Python's Generators and Iterators?](#37--what-is-the-difference-between-pythons-generators-and-iterators) - [38- What are Python documentation strings?](#38--what-are-python-documentation-strings) -- [39- Explain the use of `subn()`, `sub()`, and `split()` in the `“re”` module](#39--explain-the-use-of-subn-sub-and-split-in-the-re-module) +- [39- Explain the use of `sub()`, `subn()`, and `split()` in the `re` module](#39--explain-the-use-of-sub-subn-and-split-in-the-re-module) - [40- Define polymorphism in Python](#40--define-polymorphism-in-python) - [41- What are the differences between Wheels and Eggs?](#41--what-are-the-differences-between-wheels-and-eggs) -- [42- What is the purpose of Python non-local statements?](#42--what-is-the-purpose-of-python-non-local-statements) -- [43- How is Python exception is handled?](#43--how-is-python-exception-is-handled) +- [42- What is the purpose of the `nonlocal` statement in Python?](#42--what-is-the-purpose-of-the-nonlocal-statement-in-python) +- [43- How are exceptions handled in Python?](#43--how-are-exceptions-handled-in-python) - [44- Name the differences between functional and object-oriented programming](#44--name-the-differences-between-functional-and-object-oriented-programming) - [45- What does the `PYTHONOPTIMIZE` flag do?](#45--what-does-the-pythonoptimize-flag-do) - [46- What are descriptors? Is there a difference between a descriptor and a decorator?](#46--what-are-descriptors-is-there-a-difference-between-a-descriptor-and-a-decorator) -- [47- Generate random number](#47--generate-random-number) -- [48- What are itertools in Python?](#48--what-are-itertools-in-python) -- [49- what does itertools.islice do?](#49--what-does-itertoolsislice-do) -- [50- Why this code will never stop?](#50--why-this-code-will-never-stop) +- [47- How do you generate a random number in Python?](#47--how-do-you-generate-a-random-number-in-python) +- [48- What is `itertools` in Python?](#48--what-is-itertools-in-python) +- [49- What does `itertools.islice` do?](#49--what-does-itertoolsislice-do) +- [50- Why will this code never stop?](#50--why-will-this-code-never-stop) - [51- What is the output of this code, and why?](#51--what-is-the-output-of-this-code-and-why) -- [52- Can we chain Multiple Decorators in Python?](#52--can-we-chain-multiple-decorators-in-python) -- [53- Build a recursive function using python](#53--build-a-recursive-function-using-python) -- [54- How to implement a binary search tree using Python?](#54--how-to-implement-a-binary-search-tree-using-python) -- [55- How to implement a binary search using Python?](#55--how-to-implement-a-binary-search-using-python) -- [56- How to implement a Linked list using Python?](#56--how-to-implement-a-linked-list-using-python) -- [57- what is `collections.OrderedDict`?](#57--what-is-collectionsordereddict) -- [58- what is `collections.defaultdict`?](#58--what-is-collectionsdefaultdict) +- [52- Can we chain multiple decorators in Python?](#52--can-we-chain-multiple-decorators-in-python) +- [53- Build a recursive function using Python](#53--build-a-recursive-function-using-python) +- [54- How do you implement a binary search tree in Python?](#54--how-do-you-implement-a-binary-search-tree-in-python) +- [55- How do you implement binary search in Python?](#55--how-do-you-implement-binary-search-in-python) +- [56- How do you implement a linked list in Python?](#56--how-do-you-implement-a-linked-list-in-python) +- [57- What is `collections.OrderedDict`?](#57--what-is-collectionsordereddict) +- [58- What is `collections.defaultdict`?](#58--what-is-collectionsdefaultdict) - [59- Can we implement an `array` using Python?](#59--can-we-implement-an-array-using-python) - [60- What is the `bytes` type?](#60--what-is-the-bytes-type) -- [61- How to concatenate tuples in python?](#61--how-to-concatenate-tuples-in-python) -- [62- How to join two `sets`?](#62--how-to-join-two-sets) -- [63- What is the difference between Python's list methods append and extend?](#63--what-is-the-difference-between-pythons-list-methods-append-and-extend) -- [64- How to implement bubble sort in Python?](#64--how-to-implement-bubble-sort-in-python) -- [65- How to implement Heap sort in Python?](#65--how-to-implement-heap-sort-in-python) -- [66- How to implement Insertion sort in Python?](#66--how-to-implement-insertion-sort-in-python) -- [67- How to implement Merge sort in Python?](#67--how-to-implement-merge-sort-in-python) -- [68- How to implement Quick Sort in Python?](#68--how-to-implement-quick-sort-in-python) -- [69- How to implement Selection sort in Python?](#69--how-to-implement-selection-sort-in-python) -- [70- How to implement Shell sort in Python?](#70--how-to-implement-shell-sort-in-python) -- [71- What are the commands that are used to copy an object in Python?](#71--what-are-the-commands-that-are-used-to-copy-an-object-in-python) +- [61- How do you concatenate tuples in Python?](#61--how-do-you-concatenate-tuples-in-python) +- [62- How do you join two sets?](#62--how-do-you-join-two-sets) +- [63- What is the difference between Python's list methods `append` and `extend`?](#63--what-is-the-difference-between-pythons-list-methods-append-and-extend) +- [64- How do you implement bubble sort in Python?](#64--how-do-you-implement-bubble-sort-in-python) +- [65- How do you implement heap sort in Python?](#65--how-do-you-implement-heap-sort-in-python) +- [66- How do you implement insertion sort in Python?](#66--how-do-you-implement-insertion-sort-in-python) +- [67- How do you implement merge sort in Python?](#67--how-do-you-implement-merge-sort-in-python) +- [68- How do you implement quicksort in Python?](#68--how-do-you-implement-quicksort-in-python) +- [69- How do you implement selection sort in Python?](#69--how-do-you-implement-selection-sort-in-python) +- [70- How do you implement Shell sort in Python?](#70--how-do-you-implement-shell-sort-in-python) +- [71- How do you copy an object in Python?](#71--how-do-you-copy-an-object-in-python) - [72- What is the difference between deep and shallow copy?](#72--what-is-the-difference-between-deep-and-shallow-copy) -- [73- How can the ternary operators be used in Python?](#73--how-can-the-ternary-operators-be-used-in-python) +- [73- How can the ternary operator be used in Python?](#73--how-can-the-ternary-operator-be-used-in-python) - [74- What will be the output of the code below?](#74--what-will-be-the-output-of-the-code-below) - [75- What will be the output of the code below?](#75--what-will-be-the-output-of-the-code-below) - [76- What will be the output of the code below?](#76--what-will-be-the-output-of-the-code-below) -- [77- What is `__slots__` in python?](#77--what-is-__slots__-in-python) -- [78- What is `__contains__` in python?](#78--what-is-__contains__-in-python) +- [77- What is `__slots__` in Python?](#77--what-is-__slots__-in-python) +- [78- What is `__contains__` in Python?](#78--what-is-__contains__-in-python) - [79- What is a "callable"?](#79--what-is-a-callable) - [80- How would you `XOR` in Python?](#80--how-would-you-xor-in-python) -- [81- What is introspection/reflection and does Python support it?](#81--what-is-introspectionreflection-and-does-python-support-it) +- [81- What is introspection/reflection, and does Python support it?](#81--what-is-introspectionreflection-and-does-python-support-it) - [82- What will be the output of lines 2, 4, 6, and 8 from the following code, and why?](#82--what-will-be-the-output-of-lines-2-4-6-and-8-from-the-following-code-and-why) -- [83- Write a function that prints the least integer that is not present in a given list and cannot be represented by the summation of the sub-elements of the list](#83--write-a-function-that-prints-the-least-integer-that-is-not-present-in-a-given-list-and-cannot-be-represented-by-the-summation-of-the-sub-elements-of-the-list) +- [83- Write a function that finds the smallest positive integer that cannot be represented as the sum of any subset of a given list](#83--write-a-function-that-finds-the-smallest-positive-integer-that-cannot-be-represented-as-the-sum-of-any-subset-of-a-given-list) - [84- How do you reverse a list? Can you come up with at least three ways?](#84--how-do-you-reverse-a-list-can-you-come-up-with-at-least-three-ways) - [85- How does Python execute code?](#85--how-does-python-execute-code) - [86- What is `__pycache__`?](#86--what-is-__pycache__) -- [87- What is the unittest in Python?](#87--what-is-the-unittest-in-python) -- [88- What is the difference between xrange and range?](#88--what-is-the-difference-between-xrange-and-range) -- [89- What is the use of `//` operator in Python?](#89--what-is-the-use-of--operator-in-python) -- [90- How are dict and set implemented internally? What is the complexity of retrieving an item? How much memory do these structures consume?](#90--how-are-dict-and-set-implemented-internally-what-is-the-complexity-of-retrieving-an-item-how-much-memory-do-these-structures-consume) -- [91- What is MRO in Python? How does it work?](#91--what-is-mro-in-python-how-does-it-work) -- [92- How to distribute Python code?](#92--how-to-distribute-python-code) -- [93- How to work with Python transitive dependencies?](#93--how-to-work-with-python-transitive-dependencies) +- [87- What is `unittest` in Python?](#87--what-is-unittest-in-python) +- [88- What is the difference between `xrange` and `range`?](#88--what-is-the-difference-between-xrange-and-range) +- [89- What is the `//` operator used for in Python?](#89--what-is-the--operator-used-for-in-python) +- [90- How are `dict` and `set` implemented internally? What is the complexity of retrieving an item? How much memory do these structures consume?](#90--how-are-dict-and-set-implemented-internally-what-is-the-complexity-of-retrieving-an-item-how-much-memory-do-these-structures-consume) +- [91- What is the MRO in Python, and how does it work?](#91--what-is-the-mro-in-python-and-how-does-it-work) +- [92- How do you distribute Python code?](#92--how-do-you-distribute-python-code) +- [93- How do you manage transitive dependencies in Python?](#93--how-do-you-manage-transitive-dependencies-in-python) - [94- What is the output of this code?](#94--what-is-the-output-of-this-code) - [95- What is the output of this code?](#95--what-is-the-output-of-this-code) -- [96- What is packing and unpacking in Python?](#96--what-is-packing-and-unpacking-in-python) +- [96- What are packing and unpacking in Python?](#96--what-are-packing-and-unpacking-in-python) - [97- What's the difference between `globals()`, `locals()`, and `vars()`?](#97--whats-the-difference-between-globals-locals-and-vars) -- [98- What is the `__init__.py` module, and what is it for?](#98--what-is-the-__init__py-module-and-what-is-it-for) -- [99- How do I view object methods?](#99--how-do-i-view-object-methods) -- [100- Which is a better practice - global import or local import in Python](#100--which-is-a-better-practice---global-import-or-local-import-in-python) -- [101- what is tilde symbol `(~)` used for in Python?](#101--what-is-tilde-symbol--used-for-in-python) +- [98- What is the `__init__.py` file, and what is it for?](#98--what-is-the-__init__py-file-and-what-is-it-for) +- [99- How do you view an object's methods?](#99--how-do-you-view-an-objects-methods) +- [100- Which is better practice in Python: global imports or local imports?](#100--which-is-better-practice-in-python-global-imports-or-local-imports) +- [101- What is the tilde symbol (`~`) used for in Python?](#101--what-is-the-tilde-symbol--used-for-in-python) - [102- What is the difference between `__str__` and `__repr__`?](#102--what-is-the-difference-between-__str__-and-__repr__) -- [103- What is `lru_cache` decorator in Python?](#103--what-is-lru_cache-decorator-in-python) +- [103- What is the `lru_cache` decorator in Python?](#103--what-is-the-lru_cache-decorator-in-python) - [104- What does `__all__` do?](#104--what-does-__all__-do) - [105- List some of the dunder methods](#105--list-some-of-the-dunder-methods) -- [106- List some of the dunder methods used in Mathematical operations](#106--list-some-of-the-dunder-methods-used-in-mathematical-operations) -- [107- List some of the Dunder variables](#107--list-some-of-the-dunder-variables) -- [108- What are python frameworks for web development?](#108--what-are-python-frameworks-for-web-development) +- [106- List some of the dunder methods used in mathematical operations](#106--list-some-of-the-dunder-methods-used-in-mathematical-operations) +- [107- List some of the dunder variables](#107--list-some-of-the-dunder-variables) +- [108- What are the main Python frameworks for web development?](#108--what-are-the-main-python-frameworks-for-web-development) - [109- Write an API using Django REST](#109--write-an-api-using-django-rest) -- [110- What are the differences between Django Framework and Django REST Framework?](#110--what-are-the-differences-between-django-framework-and-django-rest-framework) -- [111- Create a `LRU Caching` using OrderedDict class](#111--create-a-lru-caching-using-ordereddict-class) -- [112- In a peaceful kingdom, there are houses numbered from 1 to n. The king announces a prize of 100 gold coins to some special group of houses. You have been given the task to determine how many sets of three houses can form a special group, where the sum of the squares of two smaller house numbers is equal to the square of the largest house number](#112--in-a-peaceful-kingdom-there-are-houses-numbered-from-1-to-n-the-king-announces-a-prize-of-100-gold-coins-to-some-special-group-of-houses-you-have-been-given-the-task-to-determine-how-many-sets-of-three-houses-can-form-a-special-group-where-the-sum-of-the-squares-of-two-smaller-house-numbers-is-equal-to-the-square-of-the-largest-house-number) +- [110- What are the differences between Django and Django REST Framework?](#110--what-are-the-differences-between-django-and-django-rest-framework) +- [111- Create an LRU cache using the `OrderedDict` class](#111--create-an-lru-cache-using-the-ordereddict-class) +- [112- In a peaceful kingdom, there are houses numbered from 1 to n. The king announces a prize of 100 gold coins for special groups of houses. Your task is to determine how many groups of three houses are special, where a group is special if the sum of the squares of the two smaller house numbers equals the square of the largest house number](#112--in-a-peaceful-kingdom-there-are-houses-numbered-from-1-to-n-the-king-announces-a-prize-of-100-gold-coins-for-special-groups-of-houses-your-task-is-to-determine-how-many-groups-of-three-houses-are-special-where-a-group-is-special-if-the-sum-of-the-squares-of-the-two-smaller-house-numbers-equals-the-square-of-the-largest-house-number) - [113- Write a solution for the Max Pairwise Product Problem](#113--write-a-solution-for-the-max-pairwise-product-problem) - [114- What is the difference between `__new__` and `__init__`? And how does object construction work?](#114--what-is-the-difference-between-__new__-and-__init__-and-how-does-object-construction-work) - [115- How is memory allocated for `list`, `tuple`, and `set`?](#115--how-is-memory-allocated-for-list-tuple-and-set) @@ -145,7 +145,7 @@ - [141- What is Pydantic, and how does it differ from `dataclasses`?](#141--what-is-pydantic-and-how-does-it-differ-from-dataclasses) - [142- Beyond basic hints — what are `Protocol`, `TypeVar`/`Generic`, and how do you actually enforce types?](#142--beyond-basic-hints--what-are-protocol-typevargeneric-and-how-do-you-actually-enforce-types) - [143- How do you find and fix a performance bottleneck in Python?](#143--how-do-you-find-and-fix-a-performance-bottleneck-in-python) -- [144- What are the security footguns every Python engineer must know?](#144--what-are-the-security-footguns-every-python-engineer-must-know) +- [144- What are the security pitfalls every Python engineer must know?](#144--what-are-the-security-pitfalls-every-python-engineer-must-know) - [145- How do you write your own context manager, and what's in `contextlib`?](#145--how-do-you-write-your-own-context-manager-and-whats-in-contextlib) - [146- What does a senior need to know about talking to a database (ORM vs Core, sessions, pooling, transactions, N+1)?](#146--what-does-a-senior-need-to-know-about-talking-to-a-database-orm-vs-core-sessions-pooling-transactions-n1) - [147- How should logging be done in a production Python service?](#147--how-should-logging-be-done-in-a-production-python-service) @@ -158,30 +158,30 @@ - [154- Generators as coroutines: `send`, `throw`, `close`, and `yield from` — and how they became `async`/`await`](#154--generators-as-coroutines-send-throw-close-and-yield-from--and-how-they-became-asyncawait) - [155- Customizing classes without a metaclass: `__init_subclass__`, `__set_name__`, and the descriptor protocol in depth](#155--customizing-classes-without-a-metaclass-__init_subclass__-__set_name__-and-the-descriptor-protocol-in-depth) -## 1- Python uses a Global Interpreter Lock. Does that mean it doesn’t use actual threads? +## 1- Python uses a Global Interpreter Lock. Does that mean it doesn't use actual threads? -No — Python threads are **real OS threads**. `threading.Thread` maps to a genuine kernel-scheduled thread (a POSIX `pthread` or a Windows thread), not a green/user-space thread. What the Global Interpreter Lock (GIL) does is narrower than "no threads": it is a single mutex that ensures only **one thread executes Python bytecode at any given moment**. The threads are real; their execution of _Python_ code is serialised. +No — Python threads are **real OS threads**. `threading.Thread` maps to a genuine kernel-scheduled thread (a POSIX `pthread` or a Windows thread), not a green/user-space thread. What the Global Interpreter Lock (GIL) does is much narrower than "no threads": it is a single mutex that ensures only **one thread executes Python bytecode at any given moment**. The threads are real; only their execution of _Python_ code is serialized (run one at a time). -The reason the GIL exists is CPython's memory management. Every object carries a reference count that is mutated constantly (on nearly every assignment, argument pass, and scope exit). Making each of those increments and decrements individually atomic — with fine-grained locks or atomic instructions — would be slow and deadlock-prone. A single interpreter-wide lock sidesteps all of that: it keeps refcounting correct, keeps single-threaded code fast, and makes writing C extensions dramatically simpler because extension authors can assume no other Python code runs concurrently unless they explicitly release the lock. +The GIL exists because of CPython's memory management. Every object carries a reference count that changes constantly (on nearly every assignment, function call, and scope exit). Making each of those increments and decrements thread-safe on its own — with fine-grained locks or atomic CPU instructions — would be slow and prone to deadlocks. A single interpreter-wide lock avoids all of that: it keeps reference counts correct, keeps single-threaded code fast, and makes C extensions much simpler to write, because extension authors can assume no other Python code runs at the same time unless they explicitly release the lock. -The practical consequence a senior engineer needs to internalise: +The practical consequences a senior engineer needs to internalize: -- **CPU-bound work does not scale across threads.** Ten threads doing heavy computation take turns on one core and finish no faster than one — often slightly slower, because of the lock-handoff overhead. The GIL is released roughly every 5 ms (tunable via `sys.setswitchinterval`) to let another thread take a turn. +- **CPU-bound work does not scale across threads.** Ten threads doing heavy computation take turns on one core and finish no faster than one thread would — often slightly slower, because of the overhead of passing the lock around. When other threads are waiting, the running thread is asked to release the GIL roughly every 5 ms (tunable via `sys.setswitchinterval`) so another thread can take a turn. - **I/O-bound work scales fine**, because a thread **releases the GIL while it waits** on a socket, disk, or subprocess. This is why threading is still the right tool for concurrent network calls or file operations. -- To get true CPU parallelism, sidestep the GIL: use `multiprocessing`/`ProcessPoolExecutor` (separate processes, each with its own interpreter and GIL), or push the hot loop into a C extension that releases the GIL (NumPy, Cython with `nogil`). +- To get true CPU parallelism, work around the GIL: use `multiprocessing`/`ProcessPoolExecutor` (separate processes, each with its own interpreter and GIL), or move the hot loop into a C extension that releases the GIL (NumPy, or Cython with `nogil`). Two clarifications that come up constantly: - **The GIL is a CPython implementation detail, not a language feature.** Jython (JVM) and IronPython (.NET) have no GIL. A very common mistake is to lump PyPy in with them — **PyPy has a GIL too.** -- **The GIL is on its way out (optionally).** PEP 703 introduced a **free-threaded build** of CPython (experimental in 3.13, `python3.13t`), which removes the GIL in favour of fine-grained locking and lets threads run Python code truly in parallel. PEP 684 added per-interpreter GILs so sub-interpreters can run concurrently. These are opt-in and still maturing; the default CPython build you deploy today still has one global GIL. +- **The GIL is becoming optional.** PEP 703 introduced a **free-threaded build** of CPython (experimental in 3.13 as `python3.13t`, and officially supported since 3.14), which removes the GIL in favor of fine-grained locking and lets threads run Python code truly in parallel. PEP 684 gave each subinterpreter its own GIL, so subinterpreters can run in parallel too. Both are opt-in and still maturing; the default CPython build you deploy today still has a single global GIL (see question 151). ## 2- Is it possible to have a producer thread reading from the network and a consumer thread writing to a file work in parallel? What about the GIL? -Yes, it is possible to have a producer thread that reads from the network and a consumer thread that writes to a file work in parallel in Python, even with the GIL in place. The GIL prevents multiple native threads from executing Python bytecodes simultaneously. It does not prevent threads from performing other operations, such as waiting for data to be available on a network socket or for a file to be written to disk. +Yes. A producer thread that reads from the network and a consumer thread that writes to a file can work in parallel in Python, even with the GIL in place. The GIL only prevents multiple threads from executing Python bytecode at the same time. It does not prevent threads from doing other work in parallel, such as waiting for data to arrive on a network socket or for a write to reach the disk. -With a producer thread reading from the network and a consumer thread writing to a file, the producer thread can block a network read operation, allowing the consumer thread to run. Similarly, the consumer thread can block a file write operation, allowing the producer thread to run. In this way, the two threads can effectively work in parallel, even though only one native thread executes Python bytecodes at a time due to the GIL. +Blocking I/O calls release the GIL while they wait. While the producer is blocked on a network read, the consumer can run; while the consumer is blocked on a file write, the producer can run. In this way, the two threads overlap their I/O and effectively work in parallel, even though only one thread executes Python bytecode at any moment. The standard way to hand data from one to the other is a thread-safe `queue.Queue`. -It is important to note that the GIL can still limit the overall performance of a program that uses multiple threads, especially if the threads are CPU-bound. In such cases, consider using an alternative implementation of Python that does not have a GIL or using a different approach to parallelism, such as the multiprocessing module or using subprocesses. +Note that the GIL can still limit the overall performance of a multithreaded program when the threads are CPU-bound. In such cases, use a different approach to parallelism, such as the `multiprocessing` module or subprocesses, or a Python build that has no GIL. ## 3- What will be the output of the following code in each step? @@ -190,95 +190,104 @@ class C: dangerous = 2 c1 = C() c2 = C() -print (c1.dangerous) +print(c1.dangerous) c1.dangerous = 3 -print (c1.dangerous) -print (c2.dangerous) +print(c1.dangerous) +print(c2.dangerous) del c1.dangerous -print (c1.dangerous) +print(c1.dangerous) C.dangerous = 3 -print (c2.dangerous) +print(c2.dangerous) ``` _The output:_ +```text +2 +3 +2 +2 +3 +``` + In this code, `C` is a class that defines a class attribute called `dangerous`. A class attribute is a variable that is shared by all instances of the class. -We create two instances of the `C` class, `c1` and `c2`, and print the value of the `dangerous` attribute for each instance. Since the dangerous attribute is a class attribute, it has the same value for `c1` and `c2`. +We create two instances of `C`, `c1` and `c2`. The first `print` shows **`2`**: `c1` has no attribute of its own called `dangerous`, so the lookup falls back to the class attribute. -Next, we set the value of the `dangerous` attribute for `c1` to **`3`**. This creates an instance attribute for `c1` that shadows the class attribute of the same name. An instance attribute is a variable specific to a particular instance of a class, and it takes precedence over any class attribute of the same name. +Next, we set `c1.dangerous` to **`3`**. This creates an _instance_ attribute on `c1` that shadows the class attribute of the same name. An instance attribute belongs to one particular object, and attribute lookup checks the instance before the class. -We then print the value of the dangerous attribute for `c1` and `c2`. The value for `c1` is **`3`** because it now has an instance attribute of that name, while the value for `c2` is still **`2`** because it only has the class attribute of that name. +We then print `dangerous` for `c1` and `c2`. The value for `c1` is **`3`** because it now has its own instance attribute, while the value for `c2` is still **`2`** because it only sees the class attribute. -Next, we delete the instance attribute for `c1` using the del statement. This removes the instance attribute, revealing the underlying class attribute of the same name. +Next, we delete the instance attribute from `c1` using the `del` statement. This removes the shadowing attribute and reveals the class attribute again, so `c1.dangerous` prints **`2`**. -Finally, we set the value of the dangerous class attribute to **`3`**. This changes the value of the class attribute for all class instances, including `c2`. When we print the value of the dangerous attribute for `c2`, it is now **`3`**. +Finally, we set the class attribute `C.dangerous` to **`3`**. This changes the value seen by every instance that does not shadow it with its own attribute, including `c2`, so the last `print` shows **`3`**. ## 4- Why are functions considered first-class objects in Python? -In Python, functions are considered first-class objects because they have the same properties as other objects in the language. Specifically, this means that functions can be: +In Python, functions are first-class objects because they can be used like any other value in the language. Specifically, functions can be: + +1. Assigned to variables and stored in data structures, just like any other object. +2. Passed as arguments to other functions. +3. Returned as values from functions. +4. Defined inside other functions. -1. Assigned to variables and stored in data structures, just like any other object -2. Passed as arguments to functions -3. Returned as values from functions -4. Defined inside other functions - The ability to treat functions as first-class objects is a powerful feature of Python that enables several functional programming patterns, such as higher-order functions, decorators, and functional programming. +This is a powerful feature that enables several functional programming patterns, such as higher-order functions, callbacks, closures, and decorators. For example, consider the following code: ```python def greet(name): - return "Hello, " + name + return "Hello, " + name greeting = greet print(greeting("John")) # prints "Hello, John" ``` -In this code, we define a function called `greet` that takes a single argument and returns a string. We then assign the function to a variable called `greeting` and call the `greeting` function just like we would call the `greet` function. This demonstrates how a function can be treated as a first-class object and assigned to a variable. +In this code, we define a function called `greet` that takes a single argument and returns a string. We then assign the function to a variable called `greeting` and call `greeting` exactly as we would call `greet`. Both names refer to the same function object. As another example, consider the following code: ```python def apply_twice(func, arg): - return func(func(arg)) + return func(func(arg)) def add_two(x): - return x + 2 + return x + 2 print(apply_twice(add_two, 10)) # prints 14 ``` -In this code, we define a function called `apply_twice` that takes another function as an argument and applies it twice to a given argument. We then define a function called `add_two` that adds two to its argument. We pass the `add_two` function to `apply_twice` as an argument, and it is used to increment the value of `10` by two twice, resulting in a final value of `14`. This demonstrates how a function can be passed as an argument to another function. +In this code, we define a function called `apply_twice` that takes another function as an argument and applies it twice to a given value. We then define `add_two`, which adds two to its argument. When we pass `add_two` to `apply_twice`, it adds two to `10` twice, giving `14`. This shows how a function can be passed as an argument to another function. ## 5- Do arguments in Python get passed by reference or value? -In Python, arguments are passed by object reference. This means that when you pass an object to a function, a reference to the object is passed rather than a copy of the object. The behavior depends on whether the object is mutable or immutable. +In Python, arguments are passed by object reference. This means that when you pass an object to a function, the function receives a reference to that object rather than a copy of it. What you observe depends on whether the object is mutable or immutable. -The function cannot modify the original object for immutable objects (e.g., numbers, strings, and tuples) because such objects cannot be changed. Instead, any operation that seems to "modify" the object creates a new object, leaving the original one unaffected. For example: +For immutable objects (e.g., numbers, strings, and tuples), the function cannot modify the original object, because such objects cannot be changed. Any operation that seems to "modify" the object actually creates a new object, leaving the original unaffected. For example: ```python def increment(x): - x += 1 + x += 1 a = 10 increment(a) print(a) # prints 10 ``` -Here, the variable `a` remains unchanged because the function increment works with a new reference to the value `11`, leaving the original value of `a` intact. +Here, `a` remains unchanged because `x += 1` creates a new integer, `11`, and rebinds only the local name `x` to it. The caller's `a` still refers to `10`. -For mutable objects (e.g., lists and dictionaries), the function operates on the original object because the reference to the same object is passed. For example: +For mutable objects (e.g., lists and dictionaries), the function works on the original object, because it received a reference to that same object. For example: ```python def append_one(lst): - lst.append(1) + lst.append(1) a = [1, 2, 3] append_one(a) print(a) # prints [1, 2, 3, 1] ``` -In this code, the `append_one` function takes an argument `lst` and appends the value `1` to the end of the list. When we pass the list `[1, 2, 3]` to the function as an argument and then print the value of `a`, the list has been modified to include the value `1` at the end. This is because the `append_one` function operates on the original list rather than a copy of the list. +In this code, the `append_one` function appends `1` to the end of the list it receives. When we pass `a` to the function and then print `a`, the list now ends with `1`, because `append_one` operated on the original list rather than on a copy. The precise term for this model is **"call by object reference"** (sometimes "call by sharing"): the parameter name inside the function is bound to the _same object_ the caller passed. This is neither C's pass-by-value (which would copy the object) nor C++'s pass-by-reference (which would let you rebind the caller's variable). The single most important distinction to be able to state is **mutation versus rebinding**: @@ -301,15 +310,15 @@ rebind(b) print(b) # [1, 2, 3] <- unchanged ``` -You can see the mechanism with `id()`: inside `mutate`, `id(lst)` equals `id(a)` (same object); inside `rebind`, the assignment makes `id(lst)` change while `id(b)` stays put. Immutable objects behave "like pass-by-value" only because you can never mutate them — you can only rebind — so there is no shared mutation to observe. +You can see the mechanism with `id()`: inside `mutate`, `id(lst)` equals `id(a)` (same object); inside `rebind`, the assignment changes `id(lst)` while `id(b)` stays the same. Immutable objects only _look_ like they are passed by value because you can never mutate them — you can only rebind — so there is no shared change to observe. -It is crucial to understand how Python passes arguments when writing functions, as it can affect the behavior of your code. If you want to modify an object that you pass to a function and have the changes persist outside the function, you must use a mutable object such as a list or a dictionary. If you want to pass an object to a function and ensure that it is not modified, you should use an immutable object such as a number, string, or tuple. +This matters when you design functions. If a function must change the caller's data in place, pass a mutable object such as a list or a dictionary and mutate it (and document that side effect); otherwise, prefer returning a new value. If you need to guarantee that a function cannot modify what you pass, pass an immutable object (such as a number, string, or tuple) or a copy. ## 6- What tools to use for linting, debugging, and profiling? **Linting and formatting** (catch problems and enforce style before the code runs): -- **Ruff** — the modern default: an extremely fast (Rust-based) linter that consolidates and replaces Flake8, isort, pydocstyle, pyupgrade, and dozens of plugins, and also ships a formatter (`ruff format`) that is a drop-in for Black. One tool, one config in `pyproject.toml`. +- **Ruff** — the modern default: an extremely fast (Rust-based) linter that consolidates and replaces Flake8, isort, pydocstyle, pyupgrade, and dozens of plugins, and also ships a formatter (`ruff format`) that is a drop-in replacement for Black. One tool, one config in `pyproject.toml`. - **Black** — the opinionated, near-zero-config formatter that ended most style debates; Ruff's formatter is compatible with it. - **Flake8 / Pylint** — the previous generation. Pylint is still valued for its deeper, more opinionated analysis (design smells, refactor hints); Flake8 (PyFlakes + pycodestyle + McCabe) is largely superseded by Ruff. - **Type checkers** are a distinct and essential category: **mypy** (the reference checker) and **Pyright** (fast, powers Pylance in VS Code) verify type annotations statically and catch a whole class of bugs before runtime. @@ -317,20 +326,20 @@ It is crucial to understand how Python passes arguments when writing functions, **Debugging:** -- **`breakpoint()`** — the built-in (Python 3.7+) that drops into the debugger at that line; it honours the `PYTHONBREAKPOINT` env var so you can redirect it to another debugger or disable it globally. This has replaced the old `import pdb; pdb.set_trace()`. -- **pdb / ipdb** — the standard library debugger (step, inspect, set breakpoints, post-mortem with `pdb.pm()`); `ipdb` adds IPython's niceties. +- **`breakpoint()`** — the built-in (Python 3.7+) that drops into the debugger at that line; it respects the `PYTHONBREAKPOINT` environment variable, so you can redirect it to another debugger or disable it globally. This has replaced the old `import pdb; pdb.set_trace()`. +- **pdb / ipdb** — the standard library debugger (step through code, inspect variables, set breakpoints, and debug a crash after the fact with `pdb.pm()`); `ipdb` adds IPython's conveniences such as tab completion and syntax highlighting. - **IDE debuggers** — PyCharm and VS Code offer graphical breakpoints, conditional breakpoints, watch expressions, and remote debugging (`debugpy`). - Humble but effective: `logging` at DEBUG level, and `rich`/`icecream` for readable inspection output. -**Profiling** (measure before you optimise — never guess): +**Profiling** (measure before you optimize — never guess): -- **cProfile** (+ `pstats`, or the `snakeviz` visualiser) — the built-in deterministic profiler for _where CPU time goes_ by function. +- **cProfile** (+ `pstats`, or the `snakeviz` visualizer) — the built-in deterministic profiler for _where CPU time goes_ by function. - **timeit** — for micro-benchmarks of small snippets, handling warm-up and repetition correctly. -- **py-spy** — a sampling profiler that attaches to a _running_ process without modifying or restarting it and emits flame graphs; the go-to for profiling production services (it replaces the abandoned Pyflame). +- **py-spy** — a sampling profiler that attaches to a _running_ process without modifying or restarting it and produces flame graphs; the go-to tool for profiling production services (it replaced the abandoned Pyflame). - **line_profiler** (`@profile`) — line-by-line timing when you need to know which _line_ in a hot function is the cost. - **Memory profiling** is a separate concern: **tracemalloc** (standard library, snapshots and diffs allocations by line), **memory_profiler** (per-line memory), and **Scalene** (profiles CPU, GPU, and memory together, and separates Python time from native/C time — often the single most informative tool). -The senior mindset behind all this: formatting and linting are automated and non-negotiable (pre-commit + CI), and optimisation is always driven by a profiler on representative data, not by intuition about what is "probably slow". +The senior mindset behind all this: formatting and linting are automated and non-negotiable (pre-commit + CI), and optimization is always driven by a profiler on representative data, not by intuition about what is "probably slow". ## 7- Give an example of filter and reduce over an iterable object @@ -356,14 +365,15 @@ In this example, we define a list of numbers and use the `filter` function to se We then use the `reduce` function from the `functools` module to compute the product of all the even numbers in the `list`. The `reduce` function takes a function and an iterable as arguments and applies the function to the elements of the iterable in a cumulative manner, returning a single result. In this case, we pass a lambda function that multiplies its arguments and pass the list of even numbers as the iterable. -Note the `list()` call around `filter`. In Python 3 `filter` returns a lazy iterator that can only be consumed once. Had we kept the raw iterator and printed it with `print(list(even_numbers))` first, that call would have exhausted it, and the later `reduce` would raise `TypeError: reduce() of empty iterable with no initial value` instead of returning `3840`. This one-shot behaviour of iterators is a common interview gotcha. +Note the `list()` call around `filter`. In Python 3, `filter` returns a lazy iterator that can only be consumed once. Had we kept the raw iterator and printed it with `print(list(even_numbers))` first, that call would have exhausted it, and the later `reduce` would raise `TypeError: reduce() of empty iterable with no initial value` instead of returning `3840`. This one-shot behavior of iterators is a common interview gotcha. -Both `filter` and `reduce` are higher-order functions, meaning they take another function as an argument. (A higher-order function is one that takes a function as an argument, returns a function, or both — `filter` returns an iterator and `reduce` returns a single accumulated value, so neither returns a new function here.) They are helpful for concisely expressing complex operations on iterable objects in Python. +Both `filter` and `reduce` are higher-order functions: functions that take another function as an argument, return a function, or both. (Here, both take a function; `filter` returns an iterator and `reduce` returns a single accumulated value.) They are useful for expressing data transformations concisely. In modern Python, though, a comprehension (`[x for x in numbers if x % 2 == 0]`) and built-ins such as `sum()` or `math.prod()` are usually preferred for readability. ## 8- What are `list` and `dict` comprehensions? -List comprehensions and dictionary comprehensions are concise ways to create new lists and dictionaries, respectively, from existing iterable objects. They are a way to transform one list (or dictionary) into another list (or dictionary) by applying a specific operation to each element in the original list. -A list comprehension consists of square brackets containing an expression followed by a `for` clause, then zero or more `for` or `if` clauses. The result is a new list computed by evaluating the expression in the context of the `for` and `if` clauses. +List comprehensions and dictionary comprehensions are concise ways to create new lists and dictionaries, respectively, from existing iterables. They build the new collection by applying an expression to each element of the source, optionally filtering elements along the way. + +A list comprehension consists of square brackets containing an expression followed by a `for` clause, then zero or more `for` or `if` clauses. The result is a new list built by evaluating the expression for every item produced by the `for` and `if` clauses. For example, suppose we have a list of numbers, and we want to create a new list that contains only the even numbers from the original list. We could do this using a list comprehension as follows: @@ -374,7 +384,7 @@ even_numbers = [x for x in numbers if x % 2 == 0] This would create a new list `even_numbers` containing only the even numbers from the original list `numbers`. -A dictionary comprehension is similar to list comprehension, but it creates a new dictionary instead of a list. It consists of a dictionary key expression followed by a `for` clause, then zero or more `for` or `if` clauses. The result is a new dictionary computed by evaluating the key and value expressions in the context of the `for` and `if` clauses. +A dictionary comprehension is similar to a list comprehension, but it creates a new dictionary instead of a list. It consists of curly braces containing a `key: value` expression pair followed by a `for` clause, then zero or more `for` or `if` clauses. The result is a new dictionary built by evaluating the key and value expressions for every item produced by those clauses. For example, suppose we have a list of strings, and we want to create a new dictionary that maps each string to its length. We could do this using a dictionary comprehension as follows: @@ -383,13 +393,13 @@ strings = ['cat', 'dog', 'bird'] lengths = {s: len(s) for s in strings} ``` -This would create a new dictionary `lengths` that maps each string to its length. +This would create a new dictionary `lengths` that maps each string to its length. Python also has set comprehensions (`{x % 3 for x in numbers}`) and generator expressions (`(x * x for x in numbers)`), which produce values lazily instead of building a collection in memory. ## 9- What do we mean when we say that a specific Lambda expression forms a closure? -A closure is a function that retains access to the variables in the environment it was defined, even after the code that defined the function has finished executing. This means that the function can still reference and modify the variables even if the function is called in a different context, such as in a different function or a different part of the program. +A closure is a function that keeps access to the variables of the scope in which it was defined, even after that enclosing function has finished executing. The function can still read those variables (and, in a regular function using `nonlocal`, rebind them) when it is called later from a completely different place in the program. -In the context of lambda expressions, a lambda expression forms a closure if it references variables from the environment in which it was defined. For example, consider the following code: +A lambda expression forms a closure when it references variables from the enclosing function in which it was defined. For example, consider the following code: ```python def make_multiplier(n): @@ -397,10 +407,9 @@ def make_multiplier(n): double = make_multiplier(2) triple = make_multiplier(3) - ``` -Here, the `make_multiplier` function returns a lambda expression that takes a single argument `x` and returns `x * n`, where `n` is the argument passed to `make_multiplier`. The lambda expression formed by `make_multiplier` is a closure because it references the variable `n` from the environment in which it was defined, even though `make_multiplier` has already returned. +Here, `make_multiplier` returns a lambda that takes a single argument `x` and returns `x * n`, where `n` is the argument passed to `make_multiplier`. The returned lambda is a closure because it references the variable `n` from its enclosing scope, even though `make_multiplier` has already returned. We can see this in action by calling the `lambda` expressions returned by `make_multiplier`: @@ -409,7 +418,7 @@ print(double(10)) # Output: 20 print(triple(10)) # Output: 30 ``` -The lambda expression returned by `make_multiplier(2)` multiplies its argument by `2`, while the lambda expression returned by `make_multiplier(3)` multiplies its argument by `3`. This is possible because the lambda expressions formed closures and retained access to the variables in the environment in which they were defined. +The lambda returned by `make_multiplier(2)` multiplies its argument by `2`, while the one returned by `make_multiplier(3)` multiplies its argument by `3`. This works because each lambda is a closure that captured its own `n`. (You can inspect the captured values through `double.__closure__`.) ## 10- Name a few differences between Python 2.x and 3.x @@ -421,10 +430,9 @@ The lambda expression returned by `make_multiplier(2)` multiplies its argument b # Python 3.x print("Hello, World!") - ``` -2. _Division operator_: In Python 2.x, the division operator (`/`) performs floor division for integers and float division for floating-point numbers. In Python 3.x, the division operator always performs float division. +2. _Division operator_: In Python 2.x, the division operator (`/`) performs floor division when both operands are integers and true (floating-point) division when either operand is a float. In Python 3.x, `/` always performs true division, and `//` is used for floor division. ```python # Python 2.x @@ -452,10 +460,10 @@ The lambda expression returned by `make_multiplier(2)` multiplies its argument b print("An exception occurred: {}".format(e)) ``` -4. _Iterators_: In Python 2.x, the `iteritems` method is used to iterate over the keys and values of a dictionary, while in Python 3.x, the `items` method is used. +4. _Iterators and views_: In Python 2.x, `dict.items()` builds a full list, so `iteritems()` is used to iterate over keys and values lazily. In Python 3.x, `items()` returns a lightweight, lazy view, and `iteritems()` no longer exists. Likewise, `range`, `map`, `filter`, and `zip` return lazy objects in Python 3.x instead of lists. ```python - # Python 2.x + # Python 2.x d = {'a': 1, 'b': 2} for key, value in d.iteritems(): print(key, value) @@ -466,7 +474,7 @@ The lambda expression returned by `make_multiplier(2)` multiplies its argument b print(key, value) ``` -5. _Unicode support_: In Python 2.x, Unicode support is not fully integrated, and the Unicode and str types are separate. In Python 3.x, Unicode is fully integrated, and the str type is used for Unicode strings. +5. _Unicode support_: In Python 2.x, `str` is a byte string, and text needs the separate `unicode` type, which makes mixing the two a common source of bugs. In Python 3.x, `str` is always Unicode text, and binary data uses the separate `bytes` type. ## 11- How is memory managed in Python? @@ -487,17 +495,17 @@ Reference counting is simple and prompt, but it has one fatal blind spot: **refe **2. The cyclic garbage collector — the backup.** To reclaim those cycles, CPython adds a second collector (the `gc` module) that periodically finds groups of container objects reachable only through cycles and frees them. Key facts: -- It only tracks **container types** (lists, dicts, sets, class instances) — the only objects that can _form_ cycles. Atomic objects like `int` and `str` are managed by refcounting alone and are never touched by the cyclic collector. -- It is **generational**: objects are grouped into three generations (0, 1, 2). New objects start in generation 0, which is collected most often; survivors are promoted to older generations that are scanned progressively less frequently. This exploits the "most objects die young" observation to keep collection cheap. -- It runs based on allocation-count thresholds (`gc.get_threshold()`), not a timer. You can trigger it (`gc.collect()`), disable it (`gc.disable()` — sometimes done in latency-sensitive or short-lived batch processes), and inspect it. +- It only tracks **container types** (lists, dicts, sets, class instances) — the only objects that can _form_ cycles. Objects that cannot hold references to other objects, like `int` and `str`, are managed by reference counting alone and are never touched by the cyclic collector. +- It is **generational**: objects are grouped into three generations (0, 1, 2). New objects start in generation 0, which is collected most often; survivors are promoted to older generations that are scanned progressively less frequently. This relies on the observation that most objects die young, which keeps collection cheap. +- It is triggered by allocation-count thresholds (`gc.get_threshold()`), not by a timer. You can run it manually (`gc.collect()`), disable it (`gc.disable()` — sometimes done in latency-sensitive services or short-lived batch jobs), and inspect it. **3. The allocator layers — where the memory actually comes from.** CPython does not call `malloc` for every object. All Python objects live in a **private heap**, fronted by a layered allocator: -- **pymalloc** is the object allocator for small objects (≤ 512 bytes — the overwhelming majority). It requests big chunks from the OS called **arenas** (256 KB), carves each arena into 4 KB **pools**, and each pool into fixed-size **blocks**. Same-sized small objects are served from the same pool, which cuts fragmentation and avoids constant OS calls. +- **pymalloc** is the object allocator for small objects (≤ 512 bytes — the overwhelming majority). It requests large chunks of memory from the OS called **arenas**, splits each arena into **pools**, and splits each pool into fixed-size **blocks** (question 153 covers the exact sizes). Small objects of the same size are served from the same pool, which reduces fragmentation and avoids constant calls to the OS. - Larger allocations bypass pymalloc and go to the system allocator. -- A crucial consequence: **freeing Python objects does not necessarily return memory to the OS.** Freed blocks go back to pymalloc's free lists for reuse; an arena is only released when _entirely_ empty. This is why a process's resident memory often stays high after a big data structure is discarded — the memory is free for reuse by Python, just not handed back. +- A crucial consequence: **freeing Python objects does not necessarily return memory to the OS.** Freed blocks go back to pymalloc's free lists for reuse; an arena is only released when _entirely_ empty. This is why a process's resident memory often stays high after a large data structure is discarded — the memory is free for Python to reuse, just not handed back to the OS. -**4. Caches that reuse objects.** CPython pre-creates and reuses certain immutable objects, so they are never really "allocated" in hot paths: **small integers −5 to 256** are singletons, and many **strings are interned** (identifiers, and short compile-time literals). This is why `256 is 256` is `True` but `257 is 257` may be `False` (see the integer-caching question). +**4. Caches that reuse objects.** CPython pre-creates and reuses certain immutable objects, so they are not re-allocated in performance-critical code: **small integers −5 to 256** are singletons, and many **strings are interned** (identifiers and short string literals known at compile time). This is why `256 is 256` is `True` but `257 is 257` may be `False` (see the integer-caching question). **5. Tools and hooks worth naming:** @@ -505,7 +513,7 @@ Reference counting is simple and prompt, but it has one fatal blind spot: **refe - `tracemalloc` — the standard-library way to trace allocations by line and diff snapshots; the right tool for finding a memory leak. - `gc.get_objects()`, `gc.get_referrers()`, and libraries like `objgraph` — for hunting down what is keeping an object alive. - `weakref` — references that _don't_ increment the count, so they don't keep an object alive (used for caches and to break cycles). -- `__del__` — a finaliser run at deallocation; unreliable for cleanup (timing isn't guaranteed, and it can even resurrect objects), so prefer context managers for releasing resources. +- `__del__` — a finalizer that runs at deallocation. It is unreliable for cleanup (its timing isn't guaranteed, and it can even bring an object back to life), so prefer context managers for releasing resources. ## 12- What will be the output of the following code? @@ -514,12 +522,11 @@ _list = ['a', 'b', 'c', 'd', 'e'] print(_list[10:]) ``` -_The output:_ -the output will be an empty list `[]`. +_The output:_ an empty list, `[]`. -The slicing syntax `list[start: end]` retrieves a subset of the elements in a list. The `start` index specifies the index of the first element to retrieve, and the `end` index specifies the element's index after the last element to retrieve. If you omit the `end` index, the slicing syntax will return all elements of the list, starting from the `start` index until the end of the list. +The slicing syntax `list[start:end]` returns a new list containing part of the original. `start` is the index of the first element to include, and `end` is exclusive: slicing stops just before it. If you omit `end`, the slice runs from `start` to the end of the list. -In this case, the list `_list` has only five elements, so the valid indices are `0` through `4`. The index `10` is out of bounds for the list, so the slicing syntax `_list[10:]` will return an empty list. +In this case, `_list` has only five elements, so the valid indices are `0` through `4`. Unlike indexing (`_list[10]` raises `IndexError`), slicing never raises an error for out-of-range bounds: they are clamped to the length of the list. So `_list[10:]` simply returns an empty list. ## 13- A palindromic number reads the same both ways. The largest palindrome made from the product of two 2-digit numbers is 9009 = 91 × 99. Find the largest palindrome made from the product of two 3-digit numbers @@ -541,37 +548,35 @@ for i in range(100, 1000): print(largest_palindrome) ``` -This code defines a function `is_palindrome` that takes a number and returns `True` if the number is a palindrome and `False` otherwise. It does this by converting the number to a string and checking if the string is equal to its reverse. +This code defines a function `is_palindrome` that returns `True` if a number is a palindrome and `False` otherwise. It does this by converting the number to a string and checking whether the string equals its reverse. -The central part of the code then iterates over all pairs of 3-digit numbers and checks the product of each pair. If the product is a palindrome and is larger than the current largest palindrome, it updates the largest palindrome. +The main part of the code then iterates over all pairs of 3-digit numbers and checks the product of each pair. If the product is a palindrome and is larger than the current largest palindrome, it becomes the new largest palindrome. -Finally, the code prints the largest palindrome. - -This code should find the largest palindrome made from the product of two 3-digit numbers. +Finally, the code prints the result: `906609` (913 × 993). Since `i * j == j * i`, you can halve the work by starting the inner loop at `i` (`range(i, 1000)`). ## 14- What is skeleton code in Python? -Skeleton code in Python is a basic set of code that provides a starting point for a new project. It typically includes a directory structure, basic configuration files, and a set of essential functions and structures that can be used as a foundation for building out the project. +Skeleton code is a minimal outline of a program or project that serves as a starting point for new work. The structure is in place — directories, configuration files, and function or class definitions — but there is little or no real logic yet. -Skeleton code is often used to provide a consistent structure and set of best practices for projects, making it easier to get started and avoid common pitfalls. It can also be used to demonstrate how to set up a basic project or to provide a starting point for learning a new programming concept or framework. +At the code level, skeleton code means _stubs_: functions and classes with their signatures and docstrings, whose bodies are just `pass`, `...`, or `raise NotImplementedError`. This lets you agree on the design and interfaces first and fill in the implementation later. -For example, a Python skeleton code might include: +At the project level, skeleton code gives every project a consistent structure and set of best practices, making it easier to get started and avoid common pitfalls. For example, a Python project skeleton might include: -- A directory structure for organizing code, tests, and documentation -- Basic configuration files, such as a `setup.py` file for packaging and distributing the project -- A `requirements.txt` file for specifying the project's dependencies -- A testing framework and sample test cases -- Documentation templates and guidelines +- A directory structure for organizing code, tests, and documentation. +- A `pyproject.toml` file with the packaging metadata, dependencies, and tool settings (older projects use `setup.py` and `requirements.txt`). +- A testing setup (for example, pytest) with sample tests. +- Tooling such as linters, pre-commit hooks, and CI configuration. +- Documentation templates, such as a README and a `docs/` folder. -Skeleton code is often created for specific types of projects, such as web applications, command-line tools, or data science projects. There are many open-source skeleton code examples available online that you can use as a starting point for your projects. +Skeletons are often tailored to specific types of projects, such as web applications, command-line tools, or data science projects. Tools such as Cookiecutter and Copier generate project skeletons from templates, and many frameworks ship their own generators (for example, `django-admin startproject`). -Using skeleton code can help you get started with a new project more quickly and can provide a set of established best practices to follow as you develop your project. However, it's important to understand that skeleton code is only a starting point, and you will need to customize and adapt it to your specific needs as you develop your project. +A skeleton helps you start a new project quickly with established practices already in place. Keep in mind, however, that it is only a starting point: you will need to adapt it to your specific needs as the project grows. -## 15- In Python classes, what is the difference between class methods and static methods? and when to use them +## 15- In Python classes, what is the difference between class methods and static methods, and when should you use each? -In Python, a class method is a method that is bound to the class and not the instance of the class. A class method can be called on the class itself, as well as on any instance of the class. A class method is defined using the `@classmethod` decorator, and it takes the class as its first argument, conventionally named `cls`. +In Python, a class method is a method that is bound to the class rather than to an instance of the class. It can be called on the class itself, as well as on any instance of the class. A class method is defined using the `@classmethod` decorator, and it receives the class as its first argument, conventionally named `cls`. -A static method is also bound to the class rather than to an instance, but unlike a class method it receives no implicit first argument at all — neither `self` nor `cls`. It behaves like a plain function that happens to live in the class namespace. A static method is defined using the `@staticmethod` decorator. +A static method is not bound to either the class or an instance: unlike a class method, it receives no implicit first argument at all — neither `self` nor `cls`. It behaves like a plain function that happens to live in the class namespace. A static method is defined using the `@staticmethod` decorator. Here's an example of how to define and use class methods and static methods in Python: @@ -604,7 +609,7 @@ result = obj.class_method("hello") result = obj.static_method("hello") ``` -In general, you should use **_class methods_** when defining a method that operates on the class itself rather than on an instance of the class. An example of this might be a factory method that creates a new class instance with some default values. You should use **_static methods_** when you need to define a method that operates on an argument or variables that are independent of the class and its instances. An example of this might be a utility function that performs some computation or transformation on its arguments but does not need to access any class or instance attributes. +In general, use **_class methods_** when a method needs the class itself rather than a particular instance. The classic example is an alternative constructor, such as `dict.fromkeys()` or `datetime.fromtimestamp()`. Because it receives `cls`, a class-method factory also builds the correct type when it is called on a subclass. Use **_static methods_** for helper functions that logically belong with the class but need neither the class nor an instance — for example, a validation or conversion function that works only on its arguments. If the helper is not closely tied to the class, a module-level function is often simpler. ## 16- Please explain the following results of the code executed on a Python shell interpreter @@ -619,58 +624,59 @@ True False ``` -_The output:_ +_Explanation:_ -This is because of the integer caching mechanism in Python. To save time and memory costs, Python always pre-loads all the small integers in the range of [-5, 256]. +This is caused by CPython's small-integer cache. To save time and memory, CPython creates every integer in the range [-5, 256] once, when the interpreter starts. Whenever your code needs an integer in that range, CPython returns a reference to the cached object instead of creating a new one. -Therefore, all the integers in [-5, 256] have been already saved in the memory. When a new integer variable in this range is declared, Python just references the cached integer to it and won’t create any new object. +So the results are explained as follows: -Therefore, the explanations of the results are: +- When `a` and `b` are assigned `256`, both names refer to the same cached object. +- When `x` and `y` are assigned `257`, each assignment creates a separate object, because 257 is outside the cached range. -- When the variables `a` and `b` were assigned to 256, they were referenced to the same memory location where the 256 was stored. They pointed to the same object. -- When the variables `x` and `y` were assigned to 257, they were two different objects in different memory locations because 257 is not on the small integers caching range. +The `is` operator checks whether two names refer to the _same object_ (identity), not whether their values are equal. Therefore, `a is b` is `True`, and `x is y` is `False`. -Since the `is` operator is to compare the memory locations of two variables, the `a is b` should output `True`, and the `x is y` should output `False`. +Keep in mind that this is a CPython implementation detail. When both assignments are compiled together (for example, `x = 257; y = 257` on one line, or inside the same function), the compiler can reuse a single constant, and `x is y` becomes `True`. The lesson is to never use `is` to compare numbers or strings; use `==` (see question 128). -## 17- In object-oriented programming, there is a concept called abstract classes. How to implement it? +## 17- In object-oriented programming, there is a concept called abstract classes. How do you implement one in Python? -In Python, an abstract class is a class that has one or more abstract methods. An abstract method is a method that has a declaration, but no implementation. Abstract methods are defined using the `abc` (abstract base class) module, which is part of the Python standard library. +In Python, an abstract class is a class that has one or more abstract methods. An abstract method is a method that the abstract class declares but that every concrete subclass must implement. Abstract classes are built with the `abc` (abstract base classes) module from the standard library. -To create an abstract class in Python, you need to do the following: +To create an abstract class in Python: -1. Import the abc module. -2. Create a class that derives from abc.ABC. +1. Import the `abc` module. +2. Create a class that inherits from `abc.ABC`. 3. Declare one or more abstract methods using the `@abc.abstractmethod` decorator. - Here is an example of an abstract class in Python: - ```python - import abc +Here is an example of an abstract class in Python: - class Animal(abc.ABC): - @abc.abstractmethod - def make_sound(self): - pass +```python +import abc - class Dog(Animal): - def make_sound(self): - print("Woof!") +class Animal(abc.ABC): + @abc.abstractmethod + def make_sound(self): + pass - class Cat(Animal): - def make_sound(self): - print("Meow!") +class Dog(Animal): + def make_sound(self): + print("Woof!") - dog = Dog() - dog.make_sound() # Output: "Woof!" +class Cat(Animal): + def make_sound(self): + print("Meow!") - cat = Cat() - cat.make_sound() # Output: "Meow!" - ``` +dog = Dog() +dog.make_sound() # Output: "Woof!" -In this example, the `Animal` class is an abstract class because it has an abstract method called `make_sound()`. The `Dog` and `Cat` classes are concrete classes because they provide an implementation for the `make_sound()` method. The `dog` and `cat` objects are instances of the `Dog` and `Cat` classes, respectively, and they can be used to call the `make_sound()` method. +cat = Cat() +cat.make_sound() # Output: "Meow!" +``` + +In this example, `Animal` is an abstract class because it has an abstract method called `make_sound()`. `Dog` and `Cat` are concrete classes because they provide an implementation of `make_sound()`. The `dog` and `cat` objects are instances of `Dog` and `Cat`, respectively, and each one uses its own version of `make_sound()`. The points that separate a textbook answer from a senior one: -- **The real enforcement is at instantiation, and it is a hard failure.** `Animal()` raises `TypeError: Can't instantiate abstract class Animal with abstract method make_sound` — and so does any subclass that _forgets_ to implement every abstract method. That is the whole value: the failure happens loudly at object-creation time, not later as a mysterious `AttributeError` at first call. +- **The real enforcement is at instantiation, and it is a hard failure.** `Animal()` raises `TypeError: Can't instantiate abstract class Animal with abstract method make_sound` — and so does any subclass that _forgets_ to implement every abstract method. That is the whole point: the error happens immediately, when the object is created, rather than later as a confusing `AttributeError` the first time the missing method is called. - **`@abstractmethod` composes with other decorators**, so you can require an abstract `property`, `classmethod`, or `staticmethod` — always with `@abstractmethod` **innermost** (closest to the method): ```python @@ -680,15 +686,15 @@ The points that separate a textbook answer from a senior one: def area(self): ... ``` -- **`register()` creates _virtual_ subclasses.** You can declare that an unrelated existing class satisfies an ABC without editing it or inheriting from it — `MyABC.register(SomeClass)` makes `issubclass`/`isinstance` return `True`. This is exactly how the `collections.abc` hierarchy (`Iterable`, `Sequence`, `Mapping`, …) recognises built-in and third-party types. +- **`register()` creates _virtual_ subclasses.** You can declare that an existing, unrelated class satisfies an ABC without editing it or making it inherit from the ABC — `MyABC.register(SomeClass)` makes `issubclass`/`isinstance` return `True`. This is exactly how the `collections.abc` hierarchy (`Iterable`, `Sequence`, `Mapping`, …) recognizes built-in and third-party types. - **Prefer inheriting `abc.ABC` over setting `metaclass=abc.ABCMeta`.** `ABC` is just a convenience base class that already uses that metaclass; you only reach for the explicit metaclass form when combining with another metaclass. -- **ABCs are nominal; `typing.Protocol` is structural.** An ABC requires you to explicitly subclass (or `register`). A `Protocol` (PEP 544) matches any object that merely _has_ the right methods — duck typing that a static type checker can verify — with no inheritance required. Reach for an ABC when you want to _share implementation_ and enforce a contract by inheritance; reach for a Protocol when you only want to describe a shape that arbitrary types can satisfy. +- **ABCs are nominal; `typing.Protocol` is structural.** In other words, an ABC matches by _name_: a class must explicitly subclass it (or be registered). A `Protocol` (PEP 544) matches by _shape_: any object that simply _has_ the right methods qualifies — duck typing that a static type checker can verify — with no inheritance required. Reach for an ABC when you want to _share implementation_ and enforce a contract by inheritance; reach for a Protocol when you only want to describe a shape that arbitrary types can satisfy. -## 18- What are `*args` and `**kwargs` in Python +## 18- What are `*args` and `**kwargs` in Python? -In Python, the `*args` and `**kwargs` syntax is used to pass a variable number of arguments to a function. +In Python, the `*args` and `**kwargs` syntax lets a function accept a variable number of arguments. -`*args` is used to pass a variable number of non-keyworded arguments to a function. It is used to pass a tuple of arguments to the function. For example: +`*args` lets a function accept any number of extra positional arguments, which it receives as a tuple. For example: ```python def my_function(arg1, *args): @@ -702,7 +708,7 @@ my_function(1, 2, 3, 4, 5) # (2, 3, 4, 5) ``` -`**kwargs` is used to pass a variable number of keyworded arguments to a function. It is used to pass a dictionary of keyword arguments to the function. For example: +`**kwargs` lets a function accept any number of extra keyword arguments, which it receives as a dictionary. For example: ```python def my_function(**kwargs): @@ -713,12 +719,12 @@ my_function(arg1=1, arg2=2, arg3=3) # Output: {'arg1': 1, 'arg2': 2, 'arg3': 3} ``` -Both `*args` and `**kwargs` are commonly used in Python to allow a function to accept a variable number of arguments. They can be useful when you want to write a function that can be flexible and handle a wide range of input parameters. +Both are useful when you want to write flexible functions that can handle a varying set of inputs. The senior-level details: - **The names are convention, not syntax.** It is the `*` and `**` that matter; `*args`/`**kwargs` are just the customary names. `*` collects extra positionals into a **tuple**, `**` collects extra keywords into a **dict**. -- **There is a fixed parameter order**, and getting it wrong is a `SyntaxError`: standard/positional parameters, then `*args`, then **keyword-only** parameters, then `**kwargs`. Anything after `*args` (or a bare `*`) can _only_ be passed by keyword: +- **There is a fixed parameter order**, and getting it wrong is a `SyntaxError`: regular (positional-or-keyword) parameters, then `*args`, then **keyword-only** parameters, then `**kwargs`. Anything after `*args` (or a bare `*`) can _only_ be passed by keyword: ```python def f(a, b, *args, key, **kwargs): ... @@ -737,17 +743,17 @@ The senior-level details: ... ``` -- **`*` and `**` also work at the _call_ site to unpack**, which is the mirror image of collecting them in the signature: `f(*my_list, **my_dict)` spreads a list into positional args and a dict into keyword args. This is why `*args`/`**kwargs` is the standard way to write a **transparent wrapper** (e.g. a decorator) that forwards whatever it received: `def wrapper(*args, **kwargs): return func(*args, **kwargs)`. +- **`*` and `**` also work at the _call_ site to unpack**, which is the mirror image of collecting them in the signature: `f(*my_list, **my_dict)` spreads a list into positional arguments and a dict into keyword arguments. This is why `*args`/`**kwargs` is the standard way to write a **transparent wrapper** (e.g., a decorator) that forwards whatever it received: `def wrapper(*args, **kwargs): return func(*args, **kwargs)`. ## 19- What is the difference between tuples, sets, and lists in Python? -In Python, tuples, sets, and lists are all data types that can be used to store collections of items. Here are the main differences between them: +In Python, tuples, sets, and lists are all built-in types for storing collections of items. Here are the main differences between them: -1. **Tuples** are immutable, which means that you cannot modify the values of the items in a tuple once it has been created. They are defined using parentheses `()` and their items are separated by commas. also, tuples are generally faster and use less memory than lists, because they do not have the overhead of the extra methods and behaviors that are associated with lists. However, the difference in performance between tuples and lists is usually small and may not be noticeable in most cases. +1. **Tuples** are ordered and immutable, which means that you cannot add, remove, or replace items once the tuple has been created. They are defined using parentheses `()`, with items separated by commas. Tuples use slightly less memory than lists and are slightly faster to create, because their size is fixed: they need no spare capacity for growth, and constant tuples can be built once at compile time. Because they are immutable, tuples are also hashable (as long as all their items are), so they can be used as dictionary keys or set members. In most programs, though, the performance difference between tuples and lists is too small to matter. -2. **Sets** are mutable, but unlike lists they do not have a specific order and do not allow duplicate items. Sets are defined using curly braces `{}` and their items are separated by commas — note that `{}` on its own creates an empty dictionary, so use `set()` for an empty set. You can add and remove items after creation with `add()`, `remove()`, and `discard()`. Sets provide much faster membership tests (`x in s`) than lists, because they are implemented using a hash table data structure, which allows for efficient insertion, deletion, and lookup of items. Their items must be hashable, which means a set can hold tuples but not lists or other sets. If you need an immutable, hashable set, use `frozenset` instead. However, sets do not maintain the order of their items, which can be a drawback if you need to preserve the order of the items in your collection. +2. **Sets** are mutable, but unlike lists they do not have a specific order and do not allow duplicate items. Sets are defined using curly braces `{}` and their items are separated by commas — note that `{}` on its own creates an empty dictionary, so use `set()` for an empty set. You can add and remove items after creation with `add()`, `remove()`, and `discard()`. Sets provide much faster membership tests (`x in s`) than lists, because they are implemented using a hash table data structure, which allows for efficient insertion, deletion, and lookup of items. Their items must be hashable, which means a set can hold tuples but not lists or other sets. If you need an immutable, hashable set, use `frozenset` instead. Because sets do not preserve insertion order, they are a poor fit when the order of items matters. -3. **Lists** are mutable, which means that you can change the values of their items after the list has been created. They are defined using square brackets `[]` and their items are separated by commas. also, Lists are generally slower and use more memory than tuples, because they are mutable and have the overhead of the extra methods and behaviors that are associated with them. However, lists are more flexible than tuples because you can modify their items after the list has been created. +3. **Lists** are ordered and mutable, which means that you can add, remove, and replace items after the list has been created. They are defined using square brackets `[]`, with items separated by commas. Lists use a little more memory than tuples because they reserve spare capacity so that appends are fast (see question 116), but this makes them far more flexible. ```python # Create a tuple @@ -774,13 +780,13 @@ s.discard(3) fs = frozenset([3, 6, 9]) ``` -Tuples are generally used when you want to store a collection of items that should not be modified, lists when you need an ordered collection that you want to be able to modify, and sets when you need fast membership tests and automatic removal of duplicates and do not care about order. +In short, use tuples for a fixed collection of items that should not change (often a record of different kinds of values, like `(name, age)`), lists for an ordered collection that you need to modify, and sets when you need fast membership tests or automatic removal of duplicates and do not care about order. ## 20- What are pickling and unpickling in Python? -In Python, "pickling" refers to the process of converting an object hierarchy (e.g., a list, dictionary, or a user-defined object) into a byte stream, and "unpickling" refers to the process of reconstructing the object hierarchy from the byte stream. +In Python, "pickling" is the process of converting an object and everything it references (e.g., a list, a dictionary, or a user-defined object) into a stream of bytes, and "unpickling" is the process of rebuilding those objects from the bytes. This is Python's built-in form of serialization. -The `pickle` module in Python provides functions for pickling and unpickling objects. To pickle an object, you can use the `pickle.dump` function, which takes the object to be pickled and a file-like object (such as a file or a byte stream) as arguments and writes the pickled object to the file-like object. To unpickle an object, you can use the `pickle.load` function, which takes a file-like object as an argument and returns the unpickled object. +The `pickle` module provides functions for pickling and unpickling objects. To pickle an object, use `pickle.dump`, which takes the object and a binary file-like object (such as a file opened in `'wb'` mode or an `io.BytesIO`) and writes the pickled bytes to it. To unpickle, use `pickle.load`, which takes a binary file-like object and returns the rebuilt object. The `pickle.dumps` and `pickle.loads` variants work directly with `bytes` instead of files. Here is an example of how to use the `pickle` module to pickle and unpickle a simple object in Python: @@ -789,7 +795,7 @@ import pickle # Define a simple object to be pickled data = {'a': [1, 2.0, 3, 4+6j], - 'b': ('string', u'Unicode string'), + 'b': ('string', 'Unicode string'), 'c': None} # Pickle the object @@ -801,13 +807,13 @@ with open('data.pkl', 'rb') as f: data_loaded = pickle.load(f) ``` -Pickling is useful for storing complex objects in a file or for sending them over a network connection. However, it is important to note that the pickle module is not intended to be secure, and it is possible to construct malicious pickle data that can execute arbitrary code when unpickled. Therefore, it is generally not recommended to use pickle to serialize and transmit sensitive data over untrusted networks or to unserialize pickle data from untrusted sources. +Pickling is useful for saving complex Python objects to a file or passing them between Python processes (this is how `multiprocessing` sends data to worker processes). However, the `pickle` module is **not secure**: a malicious pickle can execute arbitrary code when it is unpickled. Never unpickle data from an untrusted or unauthenticated source. For exchanging data with other systems, prefer a safe, language-neutral format such as JSON. ## 21- Does Python support multiple inheritance? -Yes, Python supports multiple inheritance, which means that a class can inherit from multiple superclasses (also called base classes or parent classes). Multiple inheritance can be useful when you want to define a class that inherits behavior from more than one parent class. +Yes, Python supports multiple inheritance, which means that a class can inherit from more than one superclass (also called a base class or parent class). This is useful when a class should combine behavior from several parents — for example, mixing in small, focused "mixin" classes. -To use multiple inheritance in Python, you can specify multiple superclasses in the class definition, separated by commas. For example: +To use multiple inheritance, list the superclasses in the class definition, separated by commas. For example: ```python class Base1: @@ -827,7 +833,7 @@ In this example, the `Derived` class inherits from both the `Base1` and `Base2` Note that a class body cannot be empty in Python, which is why each class above uses `pass`; a comment alone is not enough and raises an `IndentationError`. -It is important to note that Python uses a method resolution order (MRO) to determine which method should be called when a method with the same name is inherited from multiple superclasses. The MRO is computed with the C3 linearisation algorithm. It respects the order in which the superclasses are listed, but it also guarantees that a class always appears before its own parents, so it is not a simple left-to-right search. You can inspect it at any time with `ClassName.__mro__` or `ClassName.mro()`: +Python uses a method resolution order (MRO) to decide which method is called when methods with the same name are inherited from multiple superclasses. The MRO is computed with the C3 linearization algorithm. It respects the order in which the superclasses are listed, but it also guarantees that a class always appears before its own parents, so it is not a simple left-to-right search. You can inspect it at any time with `ClassName.__mro__` or `ClassName.mro()`: ```python class A: pass @@ -841,17 +847,17 @@ print([cls.__name__ for cls in D.__mro__]) Here `A` comes after both `B` and `C`, even though `B` inherits from `A`, because C3 places every class ahead of its ancestors. -For more information about multiple inheritance in Python, you can refer to the documentation on class inheritance in the Python tutorial. +See question 91 for more on how the MRO works, and question 31 for how `super()` follows it. -## 22- What are the pitfalls and problems of Python language? +## 22- What are the pitfalls and problems of the Python language? -Python's pitfalls fall into two buckets: **language-level footguns** (surprising semantics that bite everyone eventually) and **platform-level limitations** (structural properties of the runtime). A senior engineer should be able to reel off both. +Python's pitfalls fall into two groups: **language-level gotchas** (surprising behavior that catches everyone eventually) and **platform-level limitations** (built-in properties of the runtime). A senior engineer should be able to list both. **Language-level gotchas** (most have a dedicated question elsewhere in this file): - **Mutable default arguments.** `def f(x, items=[])` shares _one_ list across all calls, because defaults are evaluated once at definition time. Use `items=None` + `if items is None: items = []`. -- **Late-binding closures.** `[lambda: i for i in range(3)]` all capture the _variable_ `i`, not its value, so they all return `2`. Bind with a default arg (`lambda i=i: i`). -- **`is` vs `==`.** `is` compares identity, `==` compares value. Small-int and string caching makes `is` _appear_ to work on values (`256 is 256`) until it suddenly doesn't (`257 is 257`). Only use `is` for `None`/`True`/`False`/sentinels. +- **Late-binding closures.** The lambdas in `[lambda: i for i in range(3)]` all capture the _variable_ `i`, not its value, so they all return `2`. Capture the current value with a default argument (`lambda i=i: i`). +- **`is` vs `==`.** `is` compares identity, `==` compares value. Small-int and string caching makes `is` _appear_ to work on values (`256 is 256`) until it suddenly doesn't (`257 is 257`). Only use `is` for `None`, `True`, `False`, and sentinel objects. - **Floating-point equality.** `0.1 + 0.2 != 0.3`. Never test floats for exact equality; use `math.isclose`. - **Shared references and aliasing.** `[[0] * 3] * 2` makes two references to one inner list; `b = a` does not copy. Know shallow vs deep copy. - **Integer division and modulo sign.** `//` floors toward negative infinity, and `%` takes the sign of the divisor — different from C. @@ -860,16 +866,16 @@ Python's pitfalls fall into two buckets: **language-level footguns** (surprising **Platform-level limitations:** -- **The GIL** serialises Python bytecode execution, so threads don't give CPU-bound parallelism (use `multiprocessing` or C extensions; see the GIL questions). -- **Raw execution speed.** As a dynamically typed, interpreted language, pure-Python numeric loops are far slower than C — which is why the ecosystem pushes hot paths into NumPy/Cython/native extensions. -- **Memory footprint.** Every value is a full heap object with refcount and type-pointer overhead; a `list` of a million ints costs far more than a C array (mitigate with `array`, `__slots__`, NumPy, or generators). +- **The GIL** serializes Python bytecode execution, so threads don't give CPU-bound parallelism (use `multiprocessing` or C extensions; see the GIL questions). +- **Raw execution speed.** Because Python is dynamically typed and interpreted, pure-Python numeric loops are far slower than C — which is why the ecosystem moves performance-critical code into NumPy, Cython, or native extensions. +- **Memory footprint.** Every value is a full heap object that carries a reference count and a type pointer; a `list` of a million ints costs far more than a C array (mitigate with `array`, `__slots__`, NumPy, or generators). - **Dynamic typing defers errors to runtime.** A typo or type mismatch may only surface on the code path that hits it — the reason type hints + mypy/Pyright and good test coverage matter so much on large codebases. - **Packaging and dependency management** are historically painful (multiple competing tools, environment isolation, transitive-dependency conflicts), though `pyproject.toml`, lock files, and tools like uv/Poetry have improved it a lot. -- **Startup time and distribution.** Shipping Python to machines without an interpreter needs PyInstaller/containers; interpreter start-up is non-trivial for short-lived CLIs. +- **Startup time and distribution.** Shipping Python to machines without an interpreter requires tools like PyInstaller, or containers, and interpreter startup time is noticeable for short-lived command-line tools. -The balanced senior take: none of these make Python a poor choice — its readability, ecosystem, and development speed usually dominate. The skill is knowing _when_ a limitation actually bites (a tight CPU-bound inner loop, a memory-constrained service) and reaching for the right escape hatch, rather than avoiding Python or fighting it prematurely. +The balanced senior take: none of these make Python a poor choice — its readability, ecosystem, and development speed usually dominate. The skill is knowing _when_ a limitation actually matters (a tight CPU-bound inner loop, a memory-constrained service) and reaching for the right workaround, rather than avoiding Python or fighting it prematurely. -## 23- How to achieve multithreading in Python? +## 23- How do you achieve multithreading in Python? Multithreading in Python is achieved with the built-in `threading` module, either by instantiating `threading.Thread` with a target callable or by using `concurrent.futures.ThreadPoolExecutor` for a higher-level pool interface. @@ -897,15 +903,15 @@ with ThreadPoolExecutor(max_workers=3) as executor: results = list(executor.map(worker, range(3))) ``` -It is important to understand what this does and does not buy you. Because of the Global Interpreter Lock, only one thread executes Python bytecode at a time, so threads take rapid turns rather than truly running Python code in parallel on multiple cores. This means threading does **not** speed up CPU-bound work such as number crunching, and may even slow it down slightly because of the switching overhead. +It is important to understand what this does and does not give you. Because of the Global Interpreter Lock (on the default CPython build), only one thread executes Python bytecode at a time, so threads take rapid turns rather than truly running Python code in parallel on multiple cores. This means threading does **not** speed up CPU-bound work such as number crunching, and may even slow it down slightly because of the switching overhead. -Threads are still very effective for I/O-bound work — reading from the network, querying a database, or writing files — because a thread releases the GIL while it waits on I/O, allowing other threads to run. For CPU-bound work, use the `multiprocessing` module or `ProcessPoolExecutor` instead, which run separate processes each with their own interpreter and GIL. +Threads are still very effective for I/O-bound work — reading from the network, querying a database, or writing files — because a thread releases the GIL while it waits on I/O, allowing other threads to run. For CPU-bound work, use the `multiprocessing` module or `ProcessPoolExecutor` instead, which run separate processes, each with its own interpreter and GIL. When threads share mutable data, protect it with a `threading.Lock` or pass it through a thread-safe `queue.Queue` (see question 148). ## 24- What is the use of `with` in Python? -In Python, the with statement is used to wrap the execution of a block of code with methods defined by a context manager. A context manager is an object that defines the methods `__enter__` and `__exit__`, which are called before and after the execution of the block of code, respectively. +In Python, the `with` statement wraps the execution of a block of code with setup and cleanup logic defined by a context manager. A context manager is an object that defines the methods `__enter__` and `__exit__`, which are called before and after the block runs, respectively. -The `with` statement is used to manage resources that need to be acquired and released, such as file handles or network connections. It is particularly useful when working with resources that need to be closed or released after they are no longer needed because it ensures that the resources are properly cleaned up even if an exception is raised during the execution of the block of code. +The `with` statement is used to manage resources that must be acquired and then released, such as files, locks, network connections, or database transactions. Its key guarantee is that the cleanup code always runs — even if an exception is raised inside the block. Here is an example of how to use the `with` statement to open and read a file in Python: @@ -916,7 +922,7 @@ with open('filename.txt', 'r') as f: # the file is closed automatically here, even if f.read() raised an exception ``` -Without the `with` statement you would have to release the resource yourself in a `try`/`finally` block. The following code is equivalent to the example above, which shows what `with` saves you from writing: +Without the `with` statement, you would have to release the resource yourself in a `try`/`finally` block. The following code is equivalent to the example above and shows what `with` saves you from writing: ```python f = open('filename.txt', 'r') @@ -944,19 +950,21 @@ with managed_resource("db connection") as res: print(f"using {res['name']}") ``` +Question 145 covers writing context managers and the `contextlib` helpers in more depth. + ## 25- How are `.py`, `.pyi`, `.pyd`, and `.pyc` files different? -In Python, there are several different file types that you may encounter when working with the language: +These are the main file types you will encounter when working with Python: -1. **`.py`** files are Python source files that contain the code written in the Python language. These files can be executed by the Python interpreter, and they can be imported as modules in other Python programs. +1. **`.py`** files are Python source files. They can be run directly by the Python interpreter or imported as modules by other Python code. -2. **`.pyi`** files are Python interface files that contain type hints for Python programs. These files are used to provide type information for static type checkers, such as `mypy`, and they are not intended to be executed by the Python interpreter. +2. **`.pyi`** files are _stub_ files that contain only type hints: function and class signatures without implementations. Static type checkers such as mypy and Pyright read them to understand code that has no inline annotations (for example, a C extension or an untyped third-party library). They are never executed. -3. **`.pyd`** files (also known as Python dynamic libraries) are compiled binary files that contain compiled code written in C, C++, or other languages that can be imported and used by Python programs. These files are typically used to extend the functionality of Python by providing access to compiled code that is not written in Python. +3. **`.pyd`** files are compiled extension modules on Windows. A `.pyd` is essentially a DLL containing code written in C, C++, Rust, Cython, or a similar language, which Python can import like a normal module. On Linux and macOS, the equivalent is a `.so` file. Extension modules are used to speed up performance-critical code or to wrap existing native libraries. -4. **`.pyc`** files are compiled Python bytecode files that contain the bytecode version of Python source files. These files are not intended to be edited by hand and are usually generated automatically by the Python interpreter when a Python module is imported. +4. **`.pyc`** files contain compiled Python bytecode. The interpreter generates them automatically when a module is imported, so that later imports can skip the compilation step. They are not meant to be edited by hand. -Here is an example of how these file types may be used in a Python program: +Here is an example of how these file types relate in a Python program: ```python # foo.py @@ -968,13 +976,13 @@ import foo foo.foo() ``` -In this example, `foo.py` is a Python source file that defines a function called `foo`. `bar.py` is another Python source file that imports the `foo` module and calls the `foo` function. When `bar.py` is executed, the Python interpreter compiles the imported module to bytecode and caches it (if a current cached copy does not already exist), then executes that bytecode. In Python 3 this cache is not written next to the source as `foo.pyc`; it goes into a `__pycache__` directory with the interpreter version in the name, for example `__pycache__/foo.cpython-312.pyc`. Note that only imported modules are cached this way — the script you run directly is not. +In this example, `foo.py` is a Python source file that defines a function called `foo`. `bar.py` is another Python source file that imports the `foo` module and calls the `foo` function. When `bar.py` is executed, the Python interpreter compiles the imported module to bytecode and caches it (unless an up-to-date cached copy already exists), then executes that bytecode. In Python 3, this cache is not written next to the source as `foo.pyc`; it goes into a `__pycache__` directory with the interpreter version in the name, for example `__pycache__/foo.cpython-312.pyc`. Note that only imported modules are cached this way — the script you run directly is not. ## 26- What are decorators in Python? -Decorators are functions that are used to modify the behavior of other functions. Decorators are implemented as functions that take a function as an argument and return a modified function. They are often used to add additional functionality to an existing function, such as logging, caching, or input validation. +A decorator is a callable that takes a function and returns a replacement for it — usually a wrapper that adds behavior before or after the original function runs. Decorators are often used to add functionality such as logging, caching, access control, or input validation without changing the function's own code. -To use a decorator in Python, you define a decorator function and use the `@` symbol to specify that the function being defined is a decorator. The function being decorated is passed as an argument to the decorator function, and the decorator function returns the modified function. +To apply a decorator, write `@decorator_name` on the line above a function definition. Python passes the newly defined function to the decorator and binds the function's name to whatever the decorator returns. Here is an example of how to use a decorator in Python: @@ -999,19 +1007,19 @@ print(my_function(3, 4)) # 7 print(my_function.__name__) # 'my_function' ``` -In this example, the `my_decorator` function is a decorator that adds additional logging before and after the decorated function is called. The `my_function` function is decorated with the `my_decorator` decorator, which modifies its behavior to include the additional logging. +In this example, `my_decorator` prints a message before and after the decorated function is called. Applying it to `my_function` adds that behavior without touching the body of `my_function`. Two points are worth getting right, because they are common interview follow-ups: - **`@my_decorator` is just syntactic sugar for `my_function = my_decorator(my_function)`.** The decorator itself runs **once, at definition time**, not on every call. What runs on every call is the `wrapper` function it returned. So the decorated name is rebound to `wrapper`, and calling `my_function(3, 4)` really calls `wrapper(3, 4)`, which in turn calls the original function. -- **Always apply `functools.wraps`.** Without it the wrapper replaces the original's metadata, so `my_function.__name__` would be `'wrapper'`, the docstring would be lost, and tooling such as `help()`, debuggers, and documentation generators would report the wrong thing. `functools.wraps` copies `__name__`, `__doc__`, `__module__`, `__qualname__`, and sets `__wrapped__` so the original is still reachable. +- **Always apply `functools.wraps`.** Without it, the wrapper replaces the original's metadata, so `my_function.__name__` would be `'wrapper'`, the docstring would be lost, and tools such as `help()`, debuggers, and documentation generators would report the wrong thing. `functools.wraps` copies `__name__`, `__qualname__`, `__doc__`, `__module__`, and `__annotations__`, and sets `__wrapped__` so the original function is still reachable. -Decorators are not limited to plain functions. A decorator can also take arguments (which requires an extra level of nesting, a function returning a decorator), be implemented as a class that defines `__call__`, or be applied to a class rather than a function. The standard library uses all of these: `functools.lru_cache`, `functools.cached_property`, `dataclasses.dataclass`, and `staticmethod` / `classmethod` / `property` are all decorators. +Decorators are not limited to plain functions. A decorator can also take arguments (which requires an extra level of nesting: a function that returns a decorator), be implemented as a class that defines `__call__`, or be applied to a class rather than a function. The standard library uses all of these: `functools.lru_cache`, `functools.cached_property`, `dataclasses.dataclass`, and `staticmethod` / `classmethod` / `property` are all decorators. -## 27- How to use `self` in Python? +## 27- How do you use `self` in Python? -In Python, the `self` keyword is used to refer to the current instance of a class. It is used inside the methods of a class to access instance variables and instance methods. +In Python, `self` is the conventional name for the first parameter of an instance method. It refers to the instance the method was called on, and it is used inside methods to access that instance's attributes and other methods. Here is an example of how to use `self` in a class definition in Python: @@ -1027,11 +1035,11 @@ obj = MyClass(10) print(obj.my_method()) ``` -In this example, the `MyClass` class defines an `__init__` method that takes an argument value and assigns it to the instance variable `self.value`. The `my_method` method returns the value of self.value. +In this example, `MyClass` defines an `__init__` method that takes an argument `value` and stores it in the instance attribute `self.value`. The `my_method` method returns `self.value`. -When the `MyClass` class is instantiated with the `MyClass(10)` statement, a new instance of the class is created, and the `__init__` method is called to initialize the instance. The `obj` variable is assigned to the new instance of the class, and the `obj.my_method()` statement calls the `my_method` method on the instance, which returns the value of self.value. +When `MyClass(10)` runs, Python creates a new instance and calls `__init__` to initialize it, passing the new instance as `self`. The `obj` variable is bound to that instance, and `obj.my_method()` calls `my_method` with `obj` as `self`, so it returns `10`. -It is important to note that the `self` keyword is not a reserved word in Python and it is not required to use it in your code. However, it is a common convention in Python to use `self` to refer to the current instance of a class, and it is recommended to follow this convention when writing Python code. +Note that `self` is not a keyword or reserved word; it is only a convention. Python passes the instance as the first argument automatically — `obj.my_method()` is equivalent to `MyClass.my_method(obj)` — so you could technically give that parameter any name. However, always use `self`: other developers, linters, and IDEs all expect it. ## 28- What are namespaces in Python? @@ -1043,7 +1051,7 @@ There are several types of namespaces in Python: - **Class namespace**: Each class in Python has its own namespace, which contains the identifiers defined in the class body, such as methods and class attributes. -- **Instance namespace**: Each instance has its own namespace, held in its `__dict__`. Attributes assigned through `self` (for example `self.z = 30`) live here, not in the class namespace. This is why two instances can hold different values for the same attribute name. +- **Instance namespace**: Each instance has its own namespace, held in its `__dict__`. Attributes assigned through `self` (for example, `self.z = 30`) live here, not in the class namespace. This is why two instances can hold different values for the same attribute name. - **Function (local) namespace**: Each call to a function creates its own namespace, which contains local variables and function arguments. It is created when the function is called and destroyed when the function returns. @@ -1091,25 +1099,25 @@ print(len) # In this example, `x` is defined in the module namespace and is accessible from the global scope and from inside `foo`. `y` is local to `foo` and is not accessible outside it. `e` lives in the enclosing namespace of `inner`. `class_attr` belongs to the class namespace and is shared by every instance, whereas `z` belongs to the individual instance's namespace, which is why it appears in `obj.__dict__` rather than in `MyClass.__dict__`. -## 29- What is PEP? +## 29- What is a PEP? -PEP stands for _Python Enhancement Proposal_. PEPs are documents that describe proposed changes, improvements, and new features for Python. They are written by Python developers and are used to communicate ideas and proposals for improving the language to the Python community. +PEP stands for _Python Enhancement Proposal_. PEPs are design documents that describe proposed changes, improvements, and new features for Python, along with the reasoning behind them. They are the main way new ideas are proposed, discussed, and recorded, so the history of every major language decision can be traced through them. -There are different types of PEPs, including: +There are three types of PEPs: -- **Standards Track PEPs**: These PEPs propose changes to the Python language itself, such as new syntax or built-in functions. +- **Standards Track PEPs** propose new features or changes to Python, such as new syntax or standard library additions (for example, PEP 484, which introduced type hints). -- **Informational PEPs**: These PEPs provide information about Python-related topics, such as best practices or design patterns. +- **Informational PEPs** describe design issues or provide general guidelines without proposing a new feature (for example, PEP 20, "The Zen of Python"). -- **Process PEPs**: These PEPs describe changes to the Python development process, such as how PEPs are submitted and reviewed. +- **Process PEPs** describe changes to how Python is developed, such as how PEPs are submitted and reviewed, or propose conventions for the community (for example, PEP 8, the style guide for Python code). -PEPs are written in a standard format and are reviewed by the Python community through a process called the PEP process. The PEP process is designed to ensure that proposed changes to Python are well-documented, well-reasoned, and discussed by the community before being accepted and implemented. +Each PEP follows a standard format and goes through public review, now mostly on the Python community forum (discuss.python.org). The Python Steering Council then accepts or rejects it. This process ensures that changes to Python are well-documented, well-reasoned, and discussed by the community before they are implemented. ## 30- What are dunder methods in Python? -In Python, dunder methods (also known as "magic methods") are methods that are defined with double underscores (e.g., `__init__`, `__len__`) and are used to implement special behavior for objects. These methods are called "dunder" because they are surrounded by double underscores (i.e., "double underscore" or "dunder"). +In Python, dunder methods (also known as "magic methods" or "special methods") are methods whose names begin and end with double underscores (e.g., `__init__`, `__len__`). The name "dunder" is short for "double underscore." -Dunder methods are used to define the behavior of various built-in operations in Python, such as arithmetic operations, attribute access, and object creation and destruction. For example, the `__init__` dunder method is used to initialize an object when it is created, and the `__add__` dunder method is used to define the behavior of the `+` operator for an object, and the `__str__` dunder method is used to define the string representation of an object. +Dunder methods let your objects work with Python's built-in operations and syntax, such as arithmetic operators, attribute access, iteration, and object creation and destruction. You rarely call them directly; Python calls them for you. For example, `__init__` initializes an object when it is created, `__add__` defines the behavior of the `+` operator, and `__str__` defines the string that `str()` and `print()` produce. Here is an example of how dunder methods are used in Python: @@ -1128,16 +1136,15 @@ obj1 = MyClass(10) obj2 = MyClass(20) print(obj1 + obj2) # Calls the __add__ method print(str(obj1)) # Calls the __str__ method - ``` -In this example, the `MyClass` class defines the `__init__`, `__add__`, and `__str__` dunder methods to customize the behavior of object creation, the `+` operator, and the `str` function for instances of the class. +In this example, `MyClass` defines the `__init__`, `__add__`, and `__str__` dunder methods to customize object creation, the `+` operator, and the `str()` function for its instances. -For more information about dunder methods and how they are used in Python, you can refer to the documentation. +The complete list is in the "Data model" chapter of the Python Language Reference; questions 105, 106, and 130 cover the most commonly used ones. -## 31- What does `super` do in Python? and what is the difference between `super().__init__()` and explicit `superclass.__init__()` +## 31- What does `super` do in Python, and what is the difference between `super().__init__()` and an explicit `superclass.__init__()` call? -`super()` returns a proxy object that delegates method calls to the **next class in the method resolution order (MRO)**, starting after the current class. In the common case of single inheritance that next class is simply the parent, which is why `super()` is usually described as "referring to the parent class" — but that description is a simplification that breaks down under multiple inheritance, and the distinction is the whole point of this question. +`super()` returns a proxy object that delegates method calls to the **next class in the method resolution order (MRO)**, starting after the current class. In the common case of single inheritance, that next class is simply the parent, which is why `super()` is usually described as "referring to the parent class" — but that description is a simplification that breaks down under multiple inheritance, and the distinction is the whole point of this question. When you use `super().__init__()` in a child class, you are calling the `__init__` of whichever class comes next in the MRO of the _actual_ object being constructed. @@ -1158,7 +1165,7 @@ class Cat(Animal): cat1 = Cat("Kitty", "Siamese", "Ball") ``` -In this example, the `Cat` class is a child class of the `Animal` class. The `Cat` class has its own `__init__` method, which calls the `__init__` method of the `Animal` class using `super().__init__(name, species="Cat")`, setting the `name` and `species` attributes of the `Cat` object. +In this example, `Cat` is a child class of `Animal`. Its `__init__` method calls `Animal.__init__` through `super().__init__(name, species="Cat")`, which sets the `name` and `species` attributes on the new `Cat` object. You could instead name the parent class explicitly: @@ -1170,7 +1177,7 @@ class Cat(Animal): self.toy = toy ``` -In this single-inheritance example the two forms happen to produce the same result, but **they are not equivalent in general**. There are three real differences: +In this single-inheritance example, the two forms happen to produce the same result, but **they are not equivalent in general**. There are three real differences: **1. `super()` follows the MRO; an explicit call hard-codes one class.** With multiple inheritance, the next class in the MRO is not necessarily the class you named. Explicit calls can therefore run a shared base class more than once — the classic diamond problem: @@ -1208,13 +1215,13 @@ Note that `super().__init__()` inside `Left` calls `Right.__init__`, **not** `Ba **2. `super()` keeps the class name out of the method body.** If you rename the parent, or insert a class into the hierarchy, code using `super()` keeps working while explicit calls must be updated by hand. -**3. Explicit calls are still occasionally the right tool.** When you deliberately want one specific implementation — for example to skip a class in the MRO, or when combining classes that were never designed to cooperate — naming it explicitly is the clearer choice. +**3. Explicit calls are still occasionally the right tool.** When you deliberately want one specific implementation — for example, to skip a class in the MRO, or when combining classes that were never designed to cooperate — naming it explicitly is the clearer choice. The practical rule: use `super()` consistently throughout a hierarchy. Mixing `super()` in one class with explicit base calls in a sibling is what produces the "why did this run twice?" bugs. Note also that Python 2 required the verbose `super(Cat, self).__init__(...)`; the zero-argument `super()` shown here is Python 3 only. ## 32- What is a property decorator in Python? -In Python, the `property` decorator is a built-in function that is used to create a special kind of attribute called a "property." A property is a special kind of attribute that is defined as a method, but it is accessed like a regular attribute. +In Python, `property` is a built-in decorator that turns a method into a managed attribute called a "property." A property is defined like a method, but it is accessed like a regular attribute (without parentheses). Here is an example of how the `property` decorator is used to define a property in a class: @@ -1235,19 +1242,18 @@ class Person: self._last_name = last_name person1 = Person("John", "Doe") -print(person1.full_name) # prints "John Doe" +print(person1.full_name) # prints "John Doe" person1.full_name = "Jane Doe" -print(person1.full_name) # prints "Jane Doe" - +print(person1.full_name) # prints "Jane Doe" ``` -In this example, the `full_name` attribute is defined as a `property` using the `@property` decorator. The `full_name` property is defined as a method that returns the full name of the person, which is the combination of the `_first_name` and `_last_name` attributes. +In this example, `full_name` is defined as a property using the `@property` decorator. Its getter method returns the person's full name by combining the `_first_name` and `_last_name` attributes. The `full_name` property also has a setter, which is defined using the `@full_name.setter` decorator. The setter allows you to set the value of the `full_name` property, which in turn sets the values of the `_first_name` and `_last_name` attributes. -To use the `full_name` property, you can access it like a regular attribute, using dot notation. For example, `person1.full_name` returns the **full name** of the person, and `person1.full_name = "Jane Doe"` sets the full name of the person. +To use the `full_name` property, access it like a regular attribute, using dot notation. For example, `person1.full_name` returns the **full name** of the person, and `person1.full_name = "Jane Doe"` sets the full name of the person. -Properties are useful because they allow you to define methods that are accessed like attributes, which can make your code more readable and easier to use. They also allow you to add additional behavior to attribute access, such as data validation or type checking. +Properties are useful because they let you run code when an attribute is read or written while keeping the simple attribute syntax for callers. This is how you add behavior such as validation, type checking, or computed values to attribute access. Note that the setter above is deliberately simple; `name.split(" ")` raises a `ValueError` on a single name or a three-part name. A realistic setter validates its input, and a property can also define a **deleter**: @@ -1278,25 +1284,25 @@ print(p.full_name) # Jane Roe # p.full_name = "Cher" # ValueError: Expected 'First Last', got 'Cher' ``` -The real value of `property` is that it lets you **start with a plain attribute and add behaviour later without changing the calling code**. In languages without this feature, developers write `get_x()`/`set_x()` accessors up front just in case; in Python you write `self.x` and convert it to a property only if validation or computation becomes necessary. Callers never notice the difference. Under the hood `property` is implemented as a data descriptor, which is what allows it to intercept every read and write (see the descriptors question). +The real value of `property` is that it lets you **start with a plain attribute and add behavior later without changing the calling code**. In languages without this feature, developers write `get_x()`/`set_x()` accessors up front just in case; in Python you write `self.x` and convert it to a property only if validation or computation becomes necessary. Callers never notice the difference. Under the hood, `property` is implemented as a data descriptor, which is what allows it to intercept every read and write (see the descriptors question). ## 33- What is the difference between Cython and CPython? -Cython is a programming language that is a superset of Python, which means that it is fully compatible with Python and can be used to write Python code. Cython is designed to make it easy to write Python code that can be efficiently compiled into C or C++ code, which can then be compiled into a native machine code executable. +Cython is a programming language that is (almost entirely) a superset of Python: nearly all Python code is valid Cython, and Cython adds optional C type declarations on top. The Cython compiler translates this code into C or C++, which a C compiler then turns into a native extension module that CPython can import like any other module. -CPython, on the other hand, is the reference implementation of the Python programming language. It is written in C and is the most widely used implementation of Python. +CPython, on the other hand, is the reference implementation of the Python language — the interpreter you get from python.org. It is written in C and is by far the most widely used implementation of Python. So the two are not alternatives: Cython is a tool for producing fast extension modules, and those modules run _inside_ CPython. One of the main differences is that Cython compiles to native machine code ahead of time, whereas CPython compiles your source to **bytecode** and then executes that bytecode in a virtual machine. (CPython does not interpret the source text line by line — the `.pyc` files in `__pycache__` are the cached bytecode.) The extra step of dispatching bytecode at runtime is a large part of why pure Python is slower than C. It is worth being precise about where Cython's speed actually comes from: simply renaming a `.py` file to `.pyx` and compiling it typically yields only a modest gain, because the code still uses dynamically typed Python objects. The significant speedups come from adding static type declarations (`cdef int i`), which let Cython emit plain C operations instead of manipulating Python objects — often an order of magnitude or more on tight numeric loops. -Another difference is that Cython allows you to include C or C++ code in your Python code, which can be useful if you want to use existing C or C++ libraries or if you want to write low-level code that is not possible in pure Python. +Cython can also call C and C++ functions directly and release the GIL in sections that don't touch Python objects. This makes it a common choice for wrapping existing native libraries and for low-level code that is not possible in pure Python. -Overall, Cython is a useful tool for optimizing Python code and extending Python with C or C++ code, while CPython is the reference implementation of the Python language and is used for running Python code on most platforms. +Overall, Cython is a tool for speeding up Python code and connecting it to C or C++ code, while CPython is the standard interpreter that runs Python code on most platforms. ## 34- Specify the difference between local and global variables in Python -In Python, a local variable is a variable that is defined within a function or method and is only accessible within that function or method. A global variable is a variable that is defined outside of any function or method and is accessible from anywhere in the program. +In Python, a local variable is a variable that is assigned within a function or method and is only accessible within that function or method. A global variable is a variable that is defined at the top level of a module, outside any function or class. It is accessible from anywhere in that module, and other modules can reach it by importing the module. (Python's "global" scope is per module, not program-wide.) Here is an example of how local and global variables work in Python: @@ -1307,34 +1313,35 @@ x = 10 def some_function(): # Local variable y = 5 - print(y) # prints 5 + print(y) # prints 5 some_function() -print(x) # prints 10 -print(y) # This will cause an error because y is a local variable and is not accessible outside of the some_function() function - +print(x) # prints 10 +print(y) # NameError: y is local to some_function() and does not exist here ``` -In this example, `x` is a global variable because it is defined outside of any function or method. It is accessible from anywhere in the program, so it can be printed both inside and outside of the `some_function` function. +In this example, `x` is a global variable because it is defined outside of any function. It is accessible from anywhere in the module, so it can be printed both inside and outside of `some_function`. -`y` is a local variable because it is defined within the `some_function` function. It is only accessible within the `some_function` function and is not accessible outside of it. If you try to access `y` outside of the `some_function` function, it will cause an error because `y` is not defined in the global scope. +`y` is a local variable because it is assigned within `some_function`. It exists only while the function runs and is not accessible outside of it. Trying to access `y` outside the function raises a `NameError`, because `y` is not defined in the global scope. -It is important to note that local variables take precedence over global variables with the same name. For example: +A local variable _shadows_ a global variable with the same name. For example: ```python x = 10 def some_function(): x = 5 - print(x) # prints 5 + print(x) # prints 5 some_function() -print(x) # prints 10 +print(x) # prints 10 ``` -In this case, the `x` variable within the `some_function` function is a local variable and takes precedence over the global `x` variable. When you print `x` within the `some_function` function, it will print the value of the local `x` variable, which is **`5`**. When you print `x` outside of the function, it will print the value of the global `x` variable, which is **`10`**. +In this case, the `x` inside `some_function` is a local variable that hides the global `x`. Printing `x` inside the function shows the local value, **`5`**; printing it outside the function shows the global value, **`10`**. + +Python decides that a name is local at compile time: if a function assigns to a name _anywhere_ in its body, that name is local throughout the whole function. This is why reading `x` before assigning to it in the same function raises `UnboundLocalError` instead of reading the global. -To access the global variable from within a function, you can use the global keyword to specify that you want to access the global variable, like this: +Reading a global variable inside a function requires nothing special. To _assign_ to a global variable from within a function, declare it with the `global` keyword: ```python x = 10 @@ -1342,13 +1349,13 @@ x = 10 def some_function(): global x x = 5 - print(x) # prints 5 + print(x) # prints 5 some_function() -print(x) # prints 5 +print(x) # prints 5 ``` -In this case, the global `x` statement tells Python that you want to access the global `x` variable within the `some_function` function. This allows you to modify the value of the global `x` variable from within the function. +Here, the `global x` statement tells Python that `x` inside `some_function` refers to the module-level variable, so the assignment changes the global value. Use `global` sparingly: functions that modify module-level state are harder to test and reason about, and passing values in and returning results is usually cleaner. To assign to a variable of an _enclosing function_ instead, use `nonlocal` (see question 42). ## 35- What are Python iterators? @@ -1356,7 +1363,7 @@ In Python, an iterator is an object that produces the elements of a sequence one `__iter__` is called when iteration begins (by the `iter()` built-in, or implicitly by a `for` loop) and returns the iterator itself. `__next__` is called to retrieve the next element. When there are no more elements, `__next__` raises a `StopIteration` exception to signal that iteration is complete. -It is important to distinguish an **iterable** from an **iterator**, because interviewers ask about this constantly: +It is important to distinguish an **iterable** from an **iterator**, because interviewers ask about it often: - An **iterable** is anything you can loop over — a list, tuple, string, or dict. It implements `__iter__`, which returns a **fresh** iterator each time. Iterables can be looped over repeatedly. - An **iterator** is the object doing the actual walking. It implements both `__iter__` (returning itself) and `__next__`. An iterator is **one-shot**: once exhausted it stays exhausted, and looping over it again yields nothing. @@ -1382,17 +1389,16 @@ my_list = [1, 2, 3, 4] it = iter(my_list) # Iterate over the elements of the list -print(next(it)) # prints 1 -print(next(it)) # prints 2 -print(next(it)) # prints 3 -print(next(it)) # prints 4 -print(next(it)) # This will raise a StopIteration exception - +print(next(it)) # prints 1 +print(next(it)) # prints 2 +print(next(it)) # prints 3 +print(next(it)) # prints 4 +print(next(it)) # raises StopIteration ``` -In this example, the `iter` function is used to create an iterator object for the `my_list` list. The `next` function is then used to retrieve the elements of the list one by one. When there are no more elements to iterate over, the next function raises a _`StopIteration`_ exception. +In this example, the `iter()` function creates an iterator over `my_list`. The `next()` function then retrieves the elements one by one. When there are no more elements, `next()` raises a _`StopIteration`_ exception. -You can also use a for loop to iterate over an iterator in Python. The for loop will automatically call the `__next__` method of the iterator and will stop when a `StopIteration` exception is raised. For example: +You can also use a `for` loop to iterate over an iterator. The `for` loop calls the iterator's `__next__` method automatically and stops when `StopIteration` is raised. For example: ```python # Define a list @@ -1403,10 +1409,10 @@ it = iter(my_list) # Iterate over the elements of the list using a for-loop for i in it: - print(i) # prints 1, 2, 3, 4 + print(i) # prints 1, 2, 3, 4 ``` -Iterators are useful because they allow you to iterate over a sequence of elements in a memory-efficient way. Instead of loading the entire sequence into memory at once, an iterator loads the elements one by one as they are needed, which can save a lot of memory for large sequences. +Iterators are useful because they process elements one at a time. An iterator that computes or reads its elements on demand (such as a generator or a file object) never needs to hold the entire sequence in memory, which makes it possible to process very large, or even infinite, streams of data. You can also create your own iterators by defining the `__iter__` and `__next__` methods in a class. For example: @@ -1431,19 +1437,18 @@ it = MyIterator([1, 2, 3, 4]) # Iterate over the elements of the iterator for i in it: - print(i) # prints 1, 2, 3, 4 - + print(i) # prints 1, 2, 3, 4 ``` In this example, the `MyIterator` class defines an iterator that walks over a list of data. The `__iter__` method returns the iterator object itself, and the `__next__` method returns the next element in the sequence, raising `StopIteration` once `self.index` runs past the end of the data. -Note that `MyIterator` is its own iterator, so it inherits the one-shot behaviour described above: a second `for` loop over the same instance produces nothing, because `self.index` is already at the end. If you want an object that can be iterated many times, make it an _iterable_ instead — have `__iter__` return a new iterator (or simply be a generator function) rather than returning `self`. +Note that `MyIterator` is its own iterator, so it inherits the one-shot behavior described above: a second `for` loop over the same instance produces nothing, because `self.index` is already at the end. If you want an object that can be iterated many times, make it an _iterable_ instead — have `__iter__` return a new iterator on every call rather than returning `self`. The simplest way is to write `__iter__` itself as a generator method that uses `yield`. ## 36- What are Python generators? -In Python, a generator is a special type of function that allows you to create an iterator that generates a sequence of values on the fly. Generators are similar to iterators, but they are more memory-efficient because they do not store all of the values in memory at once. Instead, they generate the values one by one as they are needed. +In Python, a generator is the easiest way to create an iterator. A _generator function_ is a function that contains the `yield` keyword; calling it returns a _generator object_ that produces a sequence of values on the fly. Generators are iterators (see question 37), so they don't store all of their values in memory at once. Instead, they compute each value only when it is requested. -To create a generator in Python, you use the `yield` keyword instead of the `return` keyword. The `yield` keyword causes the generator to pause execution and return a value, but it does not terminate the generator function. When the generator is called again, it will resume execution from the point where it left off. +To create a generator function, use `yield` instead of `return`. Each `yield` pauses the function and hands a value to the caller, without terminating the function. When the next value is requested (with `next()` or by a `for` loop), the function resumes exactly where it left off, with all its local variables intact. Here is an example of a simple generator function in Python: @@ -1459,10 +1464,10 @@ gen = my_range(5) # Iterate over the generator for i in gen: - print(i) # prints 0, 1, 2, 3, 4 + print(i) # prints 0, 1, 2, 3, 4 ``` -In this example, the `my_range` generator function generates a sequence of numbers from `0` to `n-1`. When the generator is called, it returns an iterator object that can be used to iterate over the generated values. +In this example, the `my_range` generator function produces the numbers from `0` to `n-1`. Calling `my_range(5)` does not run any of the function body yet; it returns a generator object, and the body runs step by step as you iterate over it. You can also create a generator using a generator expression, which is a compact syntax for creating a generator. A generator expression is similar to a list comprehension, but it uses parentheses instead of square brackets and returns a generator object instead of a list. @@ -1474,11 +1479,10 @@ gen = (i for i in range(5)) # Iterate over the generator for i in gen: - print(i) # prints 0, 1, 2, 3, 4 - + print(i) # prints 0, 1, 2, 3, 4 ``` -Generators are useful when you want to generate a large sequence of values that you do not need to store in memory all at once. They allow you to generate the values one by one as they are needed, which can save a lot of memory and make your program more efficient. +Generators are useful when you need to produce a large (or even infinite) sequence of values that you do not want to store in memory all at once. Producing values one by one, only as they are needed, can save a lot of memory and lets you build efficient data-processing pipelines. ## 37- What is the difference between Python's Generators and Iterators? @@ -1488,9 +1492,7 @@ The key thing to state up front is that these are **not two competing categories - A **generator** is an iterator that Python builds for you. Calling a generator _function_ (one containing `yield`) returns a generator _object_; Python supplies `__iter__` and `__next__` automatically, and the local variables of the function body hold the state. Each `yield` suspends the function, preserving its local state, and the next `next()` call resumes exactly where it left off. -So a generator is a special case of an iterator, not an alternative to one. The practical difference is how much code you write: - -Here are the same semantics written both ways: +So a generator is a special case of an iterator, not an alternative to one. The practical difference is how much code you write. Here is the same behavior written both ways: ```python # As a hand-written iterator: a class with explicit state @@ -1523,7 +1525,7 @@ import collections.abc as abc print(isinstance(my_range(5), abc.Iterator)) # True - a generator IS an iterator ``` -Both produce identical results; the generator just replaces roughly fifteen lines of boilerplate with three. +Both produce identical results; the generator simply replaces about a dozen lines of boilerplate with a few. A point that is easy to get wrong: **both are one-shot**. Exhausting either one leaves it exhausted, and calling `iter()` on an already-consumed generator returns that same spent object rather than restarting it: @@ -1537,22 +1539,21 @@ print(it is gen) # True - iter() on an iterator returns the SAME object print(list(it)) # [] <- still exhausted, NOT restarted ``` -It is therefore wrong to say that iterators can be re-iterated while generators cannot; neither can. What _can_ be re-iterated is an **iterable** such as a list or a range, because its `__iter__` hands out a brand-new iterator on each call. If you need to traverse generated values more than once, either call the generator function again to get a fresh object, or materialise the results with `list()`. +It is therefore wrong to say that iterators can be re-iterated while generators cannot; neither can. What _can_ be re-iterated is an **iterable** such as a list or a range, because its `__iter__` hands out a brand-new iterator on each call. If you need to traverse generated values more than once, either call the generator function again to get a fresh object, or materialize the results with `list()`. -Likewise, memory efficiency is not a generator-versus-iterator distinction — both are lazy and hold only their current state. The meaningful comparison is against a **materialised sequence**: `sum(x * x for x in range(10_000_000))` holds one value at a time, while `sum([x * x for x in range(10_000_000)])` builds a ten-million-element list first. +Likewise, memory efficiency is not a generator-versus-iterator distinction — both are lazy and hold only their current state. The meaningful comparison is against a **materialized sequence**: `sum(x * x for x in range(10_000_000))` holds one value at a time, while `sum([x * x for x in range(10_000_000)])` builds a ten-million-element list first. -In short: reach for a generator by default, and write the class form only when you need behaviour a generator cannot express, such as an object that is re-iterable, introspectable, or has methods beyond iteration. +In short: reach for a generator by default, and write the class form only when you need behavior a generator cannot express, such as an object that can be iterated more than once, exposes its state for inspection, or has methods beyond iteration. ## 38- What are Python documentation strings? -In Python, documentation strings (also called docstrings) are strings that are used to document a module, class, method, or function. Docstrings are usually placed at the beginning of the code block that they document, and they are typically used to provide a brief description of what the code does and how it can be used. +In Python, documentation strings (also called docstrings) are strings that document a module, class, method, or function. A docstring is written as the first statement of the code block it documents, and it typically describes what the code does and how to use it. -In Python, docstrings are written using triple quotes (`'''` or `"""`). For example: +Docstrings are written using triple quotes (`"""` by convention, although `'''` also works). For example: ```python def some_function(arg1, arg2): - ''' - Add two numbers together. + """Add two numbers together. Parameters: arg1 (int): The first argument. @@ -1560,11 +1561,11 @@ def some_function(arg1, arg2): Returns: int: The sum of arg1 and arg2. - ''' + """ return arg1 + arg2 ``` -In this example, the docstring for the `some_function` function is a multi-line string that is placed at the beginning of the function definition. It provides a brief description of what the function does and lists the parameters and return values of the function. +In this example, the docstring of `some_function` is a multi-line string placed at the start of the function body. It gives a one-line summary of what the function does, followed by its parameters and return value. Docstrings can be accessed at runtime using the `__doc__` attribute of the object. For example: @@ -1582,7 +1583,7 @@ A few things worth knowing beyond the basics: - Docstrings differ from comments in purpose. A comment (`#`) explains implementation to someone reading the source; a docstring documents the interface and is retained at runtime for `help()`, IDEs, and documentation generators such as Sphinx. - Common formats include Google style, NumPy style, and reStructuredText. Pick one and apply it consistently. - Running Python with `-OO` **strips docstrings** from the bytecode, so avoid relying on `__doc__` for program logic. -- Docstrings in the `>>>` prompt format can be executed as tests by the `doctest` module, which keeps examples honest: +- Examples written in the `>>>` prompt format inside a docstring can be run as tests by the `doctest` module, which keeps the examples in your documentation accurate: ```python def add(a, b): @@ -1596,9 +1597,9 @@ def add(a, b): # python -m doctest yourfile.py -> runs the example and checks the output ``` -## 39- Explain the use of `subn()`, `sub()`, and `split()` in the `“re”` module +## 39- Explain the use of `sub()`, `subn()`, and `split()` in the `re` module -`re` is the Python module for regular expression matching. Among its functions (not modules — they are all functions inside the single `re` module) are three used for editing strings: `sub()`, `subn()`, and `split()`. +`re` is Python's standard module for regular expressions. Three of its functions are used for editing strings: `sub()`, `subn()`, and `split()`. 1. **`sub(pattern, repl, string)`**: Finds every non-overlapping match of `pattern` and replaces it with `repl`, returning the resulting **string**. 2. **`subn(pattern, repl, string)`**: Does exactly the same substitution, but returns a **tuple** of `(new_string, number_of_substitutions)` — not just the count. @@ -1633,23 +1634,23 @@ print(re.split(r"(\d+)", "a1b2c")) # ['a', '1', 'b', '2', 'c'] ``` -A practical note: if you use the same pattern repeatedly, compile it once with `re.compile()` and call the methods on the resulting pattern object. `re` does cache compiled patterns internally, but compiling explicitly is clearer and avoids re-parsing the pattern string on every call. +A practical note: if you use the same pattern repeatedly, compile it once with `re.compile()` and call the methods on the resulting pattern object. `re` does cache compiled patterns internally, but compiling explicitly is clearer, skips the cache lookup on every call, and doesn't depend on the cache's limited size. ## 40- Define polymorphism in Python -In Python, polymorphism refers to the ability of a function or method to behave differently depending on the data type of the arguments passed to it. +Polymorphism is the ability to use a single interface — the same function call or method name — with objects of different types, where each type provides its own behavior. -Polymorphism is a key feature of object-oriented programming (OOP) and allows you to write code that is more flexible and reusable. It allows you to define a function or method that can accept different types of arguments and perform different actions based on the type of arguments. +Polymorphism is a key feature of object-oriented programming (OOP). Code written against a shared interface works with any object that supports it, which makes that code more flexible and reusable. There are two main ways to implement polymorphism in Python: -1. **Method overriding**: Defining a method in a subclass with the same name as one in the superclass, but with different behaviour. This is the most common form of polymorphism in Python. +1. **Method overriding**: Defining a method in a subclass with the same name as one in the superclass, but with different behavior. This is the most common form of polymorphism in Python. 2. **Duck typing**: The most Pythonic form. Python does not require a shared base class at all — if an object provides the method being called, it can be used. "If it walks like a duck and quacks like a duck, it is a duck." What matters is the interface an object supports at runtime, not its position in a class hierarchy. Python does **not** support method overloading in the C++/Java sense: defining two methods with the same name simply means the second definition replaces the first. The same flexibility is achieved with default arguments, `*args`/`**kwargs`, or `functools.singledispatch` for type-based dispatch. -Here is an example of polymorphism using method overriding. Note that the point of polymorphism is calling the _same_ method on _different_ types and getting type-appropriate behaviour, so the loop below is the part that matters: +Here is an example of polymorphism using method overriding. Note that the point of polymorphism is calling the _same_ method on _different_ types and getting type-appropriate behavior, so the loop below is the part that matters: ```python import math @@ -1700,11 +1701,11 @@ for shape in [Rectangle(10, 20), Circle(5), Triangle(6, 4)]: # Triangle: 12.00 ``` -The base class is still useful as documentation and to fail loudly when a subclass forgets to implement `area()` — using `abc.ABC` with `@abstractmethod` makes that failure happen at instantiation time rather than at first call. Built-in polymorphism works the same way: `len()` works on lists, strings, and dicts because each type implements `__len__`. +The base class is still useful as documentation and to raise a clear error when a subclass forgets to implement `area()` — using `abc.ABC` with `@abstractmethod` makes that failure happen at instantiation time rather than at first call. Built-in polymorphism works the same way: `len()` works on lists, strings, and dicts because each type implements `__len__`. ## 41- What are the differences between Wheels and Eggs? -In Python, wheels and eggs are two different types of distribution formats for Python packages. +Wheels and eggs are two formats for distributing built Python packages. A **wheel** (`.whl`) is the modern standard built-distribution format, defined by **PEP 427**. It is a ZIP archive containing the package files plus a `.dist-info` metadata directory, laid out so that installing amounts to unpacking files into place. Wheels are produced today with `python -m build` (historically `setup.py bdist_wheel`) and installed with `pip`. @@ -1712,19 +1713,19 @@ An **egg** (`.egg`) is the older, `setuptools`-specific format that predates any The key differences: -- **Standardisation**: The wheel format is a documented interoperable standard (PEP 427). The egg format was never standardised — it was whatever `setuptools` happened to implement. +- **Standardization**: The wheel format is a documented interoperable standard (PEP 427). The egg format was never standardized — it was whatever `setuptools` happened to implement. - **Code execution at install time**: This is the most important practical difference. Installing an egg could execute arbitrary code by running the package's `setup.py`. Installing a wheel does not run any project code; `pip` simply unpacks the archive and moves files into place. That makes installation faster, reproducible, and safer. - **Importable vs. install-only**: Eggs were designed to be importable directly — an `.egg` file could be placed on `sys.path` and imported without being unpacked, and eggs carried runtime machinery such as `pkg_resources` entry points. A wheel is purely a _distribution_ format: it is never imported directly, only installed. -- **Build isolation and compiled code**: Wheels encode the target Python version, ABI, and platform in the filename (for example `numpy-1.26.0-cp312-cp312-win_amd64.whl`), so `pip` can pick the correct prebuilt binary and skip compiling C extensions from source. This is why installing packages such as NumPy or Pandas is now near-instant rather than requiring a compiler. +- **Prebuilt binaries for compiled code**: Wheels encode the target Python version, ABI (the binary interface of the interpreter), and platform in the filename (for example, `numpy-1.26.0-cp312-cp312-win_amd64.whl`), so `pip` can pick the correct prebuilt binary and skip compiling C extensions from source. This is why installing packages such as NumPy or pandas now takes seconds and needs no compiler. - **Tooling support**: `pip` installs wheels. `easy_install`, the only installer for eggs, no longer exists in current `setuptools`. In summary, wheels are the standard and only recommended distribution format; eggs are obsolete and should not be produced for new projects. You may still encounter `.egg-info` directories or `pkg_resources` in older codebases, but new packaging should target wheels — and `importlib.metadata` has replaced `pkg_resources` for reading package metadata at runtime. -## 42- What is the purpose of Python non-local statements? +## 42- What is the purpose of the `nonlocal` statement in Python? The `nonlocal` statement lets an inner function **rebind** a variable that belongs to an enclosing (but not global) scope. Without it, any assignment inside a function creates a brand-new local variable, shadowing the outer one and leaving the original untouched. @@ -1766,13 +1767,13 @@ broken_counter()() `nonlocal` differs from `global` in what it targets: `global` rebinds a name in the module namespace, while `nonlocal` rebinds a name in the nearest enclosing _function_ scope. `nonlocal` also requires the name to already exist in an enclosing scope — if it does not, you get a `SyntaxError` at compile time rather than silently creating one. -This matters most for closures that accumulate state, such as counters, memoisation caches, and accumulator callbacks. That said, a class or a `functools` helper is often clearer than a closure that mutates captured state, so reach for `nonlocal` when it genuinely simplifies the code. +This matters most for closures that accumulate state, such as counters, memoization caches, and accumulator callbacks. That said, a class or a `functools` helper is often clearer than a closure that mutates captured state, so use `nonlocal` only when it genuinely simplifies the code. -## 43- How is Python exception is handled? +## 43- How are exceptions handled in Python? -In Python, exceptions are handled using the `try` and except statements. +In Python, exceptions are handled using `try` and `except` statements. -Here's an example of how you can use the `try` and except statements to handle an exception: +Here's an example of how you can use `try` and `except` to handle an exception: ```python try: @@ -1783,7 +1784,7 @@ except ValueError: print('Invalid input') ``` -In this example, the `try` block contains code that might cause a _ValueError_ exception to be raised (in this case, attempting to convert the string 'foo' to an integer). If the exception is raised, the execution of the `try` block is halted, and control is transferred to the `except` block. The `except` block contains code that is executed to handle the exception. In this case, it prints an error message to the console. +In this example, the `try` block contains code that might raise a `ValueError` (here, converting the string `'foo'` to an integer). When the exception is raised, the rest of the `try` block is skipped, and control jumps to the matching `except` block, which handles the exception. In this case, it prints an error message. You can also specify multiple `except` blocks to handle different types of exceptions: @@ -1797,10 +1798,9 @@ except ValueError: except TypeError: # Code to handle the TypeError exception goes here print('Invalid type') - ``` -You can also use the `else` clause to specify a block of code that should be executed if no exceptions are raised in the try block: +You can also use the `else` clause to specify a block of code that should run only if no exception is raised in the `try` block: ```python try: @@ -1812,12 +1812,11 @@ except ValueError: else: # Code to be executed if no exceptions are raised goes here print(x) - ``` -In this example, the `else` block will be executed if the `try` block completes successfully (i.e. if no exceptions are raised), and it will print the value of `x` to the console. +In this example, the `else` block runs if the `try` block completes successfully (i.e., if no exception is raised), and it prints the value of `x`. -Finally, you can use the `finally` clause to specify a block of code that should always be executed, regardless of whether an exception is raised or not: +Finally, you can use the `finally` clause to specify a block of code that always runs, whether or not an exception is raised: ```python try: @@ -1831,7 +1830,7 @@ finally: print('Done') ``` -In this example, the `finally` block will be executed after the `try` block, regardless of whether an exception is raised or not. It will print the message **`'Done'`** to the console. `finally` runs even if the exception is _not_ caught and propagates upward, and even if the `try` block exits early via `return`, `break`, or `continue` — which is what makes it reliable for cleanup. +In this example, the `finally` block runs after the `try` block (and after the `except` block, if one ran), whether or not an exception was raised. It prints **`'Done'`**. `finally` runs even if the exception is _not_ caught and propagates upward, and even if the `try` block exits early via `return`, `break`, or `continue` — which is what makes it reliable for cleanup. Putting the four clauses together, the full form reads: @@ -1885,7 +1884,7 @@ except InsufficientFundsError as e: Custom exceptions should subclass `Exception` (not `BaseException`, which also covers `KeyboardInterrupt` and `SystemExit` — signals you almost never want to intercept). By convention (PEP 8) name them with an **`Error` suffix**, not `Exception` — `InsufficientFundsError`, not `InsufficientFundsException` (a habit worth unlearning if you come from Java). Whatever arguments you pass to an exception's constructor are stored on its **`.args`** tuple, so `e.args[0]` retrieves the original message even when you didn't define custom attributes. -**Re-raising and chaining.** A bare `raise` inside a handler re-raises the current exception with its original traceback intact, which is the right way to log and pass along. `raise ... from e` records the original cause: +**Re-raising and chaining.** A bare `raise` inside a handler re-raises the current exception with its original traceback intact, which is the right way to log an error and let it continue to propagate. `raise ... from e` records the original cause: ```python try: @@ -1900,10 +1899,10 @@ except FileNotFoundError as e: - `except:` or `except Exception:` with an empty or `pass` body — this hides real bugs and makes failures silent. Catch only what you can actually handle. - A bare `except:` also catches `KeyboardInterrupt` and `SystemExit`, making a program impossible to interrupt with Ctrl-C. Use `except Exception:` if you really must be broad. -- **Returning from a `finally` block.** A `return` (or `break`/`continue`) inside `finally` overrides any `return` value — or any in-flight exception — coming from the `try`/`except`, silently swallowing errors. Keep `finally` for cleanup only. +- **Returning from a `finally` block.** A `return` (or `break`/`continue`) inside `finally` overrides any `return` value — or any exception still being raised — from the `try`/`except`, silently swallowing errors. Keep `finally` for cleanup only. (Since Python 3.14, the compiler emits a `SyntaxWarning` for this, per PEP 765.) - Using exceptions for ordinary control flow where a simple conditional is clearer. -That said, Python idiom favours **EAFP** — "easier to ask forgiveness than permission" — over defensive pre-checks. Attempting the operation and handling the exception is usually preferred to checking first, since the check can race or miss cases: +That said, Python idiom favors **EAFP** — "easier to ask forgiveness than permission" — over defensive pre-checks. Attempting the operation and handling the exception is usually preferred to checking first, because the situation can change between the check and the operation (a race condition), and a check can miss cases: ```python # EAFP - idiomatic Python @@ -1921,19 +1920,21 @@ else: ## 44- Name the differences between functional and object-oriented programming -Functional programming and object-oriented programming are two programming paradigms that are commonly used in Python. Each paradigm has its own set of characteristics and approaches to solving problems, and they can be used in different situations depending on the needs of the project. +Functional programming and object-oriented programming are two programming paradigms (styles of structuring code), and Python supports both. Each has its own approach to solving problems, and each fits different situations depending on the needs of the project. Here are some key differences between functional programming and object-oriented programming in Python: 1. **Data model**: In functional programming, data is treated as immutable and functions are used to transform data. In object-oriented programming, data is encapsulated in objects and accessed through methods. -2. **State**: In functional programming, the state is typically avoided or minimized, and functions are designed to be pure and side-effect-free. In object-oriented programming, objects have an internal state that can be modified through methods. +2. **State**: In functional programming, mutable state is avoided or minimized, and functions are designed to be _pure_: their output depends only on their inputs, and they have no side effects. In object-oriented programming, objects have internal state that can be modified through methods. 3. **Inheritance**: In object-oriented programming, inheritance is used to create a hierarchy of classes and reuse code between classes. In functional programming, inheritance is not typically used, and functions are composed and combined to create new functionality. 4. **Polymorphism**: In object-oriented programming, polymorphism is usually achieved by overriding methods in subclasses, so the same call dispatches to different implementations depending on the object's type. In functional programming, the same flexibility comes from higher-order functions and generic functions that work across any type supporting the required operations — in Python, `functools.singledispatch` provides exactly this type-based dispatch without a class hierarchy. (Currying, sometimes cited here, is about partial application of arguments rather than polymorphism; `functools.partial` is Python's version.) -5. **Concurrency**: In functional programming, concurrency is typically easier to achieve because functions are pure and do not depend on state. In object-oriented programming, concurrency can be more challenging because objects have an internal state that can be modified concurrently. +5. **Concurrency**: In functional programming, concurrency is typically easier because pure functions share no mutable state, so there is nothing to protect with locks. In object-oriented programming, concurrency can be more challenging because several threads may modify the same object's state at the same time. + +In practice, idiomatic Python mixes both: classes to model things that have state and behavior, and pure functions, comprehensions, and generators for transforming data. ## 45- What does the `PYTHONOPTIMIZE` flag do? @@ -1951,7 +1952,7 @@ Despite the name, it performs very little optimization. It does exactly three th - Everything from level 1, plus **docstrings are stripped** from the compiled bytecode (`__doc__` becomes `None`). -That is the complete list. It does **not** inline functions, specialise calls, or eliminate dead code in any general sense. Optimisations such as constant folding and peephole optimisation are performed by the compiler on every run regardless of this setting, so they are not something `-O` turns on. +That is the complete list. It does **not** inline functions, specialize calls, or eliminate dead code in any general sense. Optimizations such as constant folding and peephole optimization are performed by the compiler on every run regardless of this setting, so they are not something `-O` turns on. ```bash python -O script.py # strip asserts, __debug__ = False @@ -1972,18 +1973,18 @@ def f(): print(__debug__) # True normally, False under -O print(f.__doc__) # 'A docstring.' normally, None under -OO -assert False, "boom" # raises AssertionError normally, silently skipped under -O +assert False, "boom" # raises AssertionError normally, silently skipped under -O ``` -The practical caution is the opposite of what is often assumed. This flag is aimed at production use — trimming assertion overhead and shrinking bytecode — not at debugging or profiling. The danger is that **stripping asserts changes behaviour if any assert is doing real work.** Code such as `assert user.is_authenticated` or an assert with a side effect silently stops running under `-O`. Assertions are for catching programmer errors during development; never use them to validate user input, enforce permissions, or perform any check your program depends on at runtime — raise a real exception instead. Similarly, `-OO` breaks any library that inspects `__doc__` at runtime, such as some CLI frameworks and doctest. +The practical caution is the opposite of what many people assume. This flag is aimed at production use — removing the cost of assertions and shrinking bytecode — not at debugging or profiling. The danger is that **stripping asserts changes behavior if any assert is doing real work.** Code such as `assert user.is_authenticated` or an assert with a side effect silently stops running under `-O`. Assertions are for catching programmer errors during development; never use them to validate user input, enforce permissions, or perform any check your program depends on at runtime — raise a real exception instead. Similarly, `-OO` breaks any library that inspects `__doc__` at runtime, such as some CLI frameworks and doctest. -The speed benefit is marginal in most programs. If performance is the real goal, profile first, then consider algorithmic fixes, optimised libraries, or an alternative runtime such as PyPy — all of which will matter far more than this flag. +The speed benefit is marginal in most programs. If performance is the real goal, profile first, then consider algorithmic fixes, optimized libraries, or an alternative runtime such as PyPy — all of which will matter far more than this flag. ## 46- What are descriptors? Is there a difference between a descriptor and a decorator? -In Python, a descriptor is an object attribute with "binding behavior", which means that it has the ability to define how it is accessed and set. Descriptors are implemented using a set of special methods, known as the descriptor protocol, which consists of the `__get__`, `__set__`, and `__delete__` methods. +In Python, a descriptor is an object, stored as a class attribute, that controls what happens when that attribute is read, written, or deleted through an instance. It does this by implementing one or more methods of the _descriptor protocol_: `__get__`, `__set__`, and `__delete__`. -Descriptors are a way to define custom attribute access behavior in Python. They can be used to implement properties, methods, or any other attribute type with custom behavior. For example, you might use a descriptor to implement a lazy evaluation of an attribute or to provide read-only access to an attribute. +Descriptors let you define custom attribute-access behavior once and reuse it across attributes and classes. They can be used to implement properties, methods, or any other attribute with custom behavior. For example, you might use a descriptor to compute an attribute lazily (only on first access), to validate values on assignment, or to make an attribute read-only. Descriptors come in two kinds, and the difference determines lookup precedence: @@ -2055,27 +2056,27 @@ print(r.width) # 3 # Rect(-1, 4) # ValueError: width must be positive, got -1 ``` -A **decorator**, by contrast, is a callable that takes a function (or class) and returns a replacement, applied with the `@` syntax to extend behaviour without modifying the original source. +A **decorator**, by contrast, is a callable that takes a function (or class) and returns a replacement, applied with the `@` syntax to extend behavior without modifying the original source. -There is a difference between a descriptor and a decorator in Python. A descriptor is an object attribute with binding behavior, whereas a decorator is a function that takes another function and extends its behavior. Descriptors are implemented using the descriptor protocol, which consists of the `__get__`, `__set__`, and `__delete__` methods, whereas decorators are implemented as functions that take a function as an argument and return a modified version of the function. +So yes, they are different things. A descriptor customizes _attribute access_ and does its work every time the attribute is accessed; it is defined by the `__get__`/`__set__`/`__delete__` protocol. A decorator transforms a function or class _once, at definition time_; it is simply a callable that takes an object and returns a replacement. The two are often used together, which is where the confusion comes from: `@property` uses decorator _syntax_ to create a descriptor _object_. -## 47- Generate random number +## 47- How do you generate a random number in Python? You can generate a random number using the `random` module. ```python import random -random.randint(1, 100) # random integer, 1 and 100 both included -random.randrange(0, 100, 2) # random even integer in [0, 100) -random.random() # random float in [0.0, 1.0), e.g. 0.3366241606464734 -random.uniform(1.0, 10.0) # random float in [1.0, 10.0] -random.choice([1, 2, 3]) # a random element from a sequence -random.sample(range(100), 5)# 5 distinct elements, without replacement -random.shuffle(my_list) # shuffles a list in place, returns None +random.randint(1, 100) # random integer, 1 and 100 both included +random.randrange(0, 100, 2) # random even integer in [0, 100) +random.random() # random float in [0.0, 1.0), e.g. 0.3366241606464734 +random.uniform(1.0, 10.0) # random float in [1.0, 10.0] +random.choice([1, 2, 3]) # a random element from a sequence +random.sample(range(100), 5) # 5 distinct elements, without replacement +random.shuffle(my_list) # shuffles a list in place, returns None ``` -You can set a seed with `random.seed(x)`, where `x` is the value used to initialise the generator. Seeding makes the sequence **reproducible**, which is what you want in tests and simulations: +You can set a seed with `random.seed(x)`, where `x` is the value used to initialize the generator. Seeding makes the sequence **reproducible**, which is what you want in tests and simulations: ```python import random @@ -2087,9 +2088,9 @@ random.seed(42) # same seed... print([random.randint(1, 10) for _ in range(3)]) # ...same sequence ``` -Seeding with the current time (`random.seed(time.time())`) is unnecessary — the generator already seeds itself from the operating system's entropy at import time, so you only call `seed()` when you specifically want reproducibility. +Seeding with the current time (`random.seed(time.time())`) is unnecessary — the generator already seeds itself from the operating system's source of randomness at import time, so you only call `seed()` when you specifically want reproducibility. -One important caveat worth raising in an interview: **`random` is not cryptographically secure.** It uses a Mersenne Twister, which is fast and statistically excellent but fully predictable — an observer who sees enough output can recover the internal state and predict all future values. For passwords, tokens, session IDs, or anything security-related, use the `secrets` module instead: +One important caveat worth raising in an interview: **`random` is not cryptographically secure.** It uses the Mersenne Twister algorithm, which is fast and statistically excellent but fully predictable — anyone who sees enough of its output can reconstruct its internal state and predict all future values. For passwords, tokens, session IDs, or anything security-related, use the `secrets` module instead: ```python import secrets @@ -2099,13 +2100,11 @@ secrets.token_hex(16) # secure random hex string, e.g. for a token secrets.choice(['a', 'b']) # secure choice from a sequence ``` -## 48- What are itertools in Python? +## 48- What is `itertools` in Python? -The `itertools` module is a Python module that provides a number of functions that are helpful when working with iterators. Iterators are objects that allow you to iterate over a sequence of values, such as a list or a string. +The `itertools` module is part of the standard library and provides fast, memory-efficient building blocks for working with iterators (objects that produce a sequence of values one at a time; see question 35). -Here are a few examples of functions that are available in the `itertools` module: - -Everything in `itertools` returns a **lazy iterator**, so values are produced on demand rather than built up in a list. That is what makes it usable with infinite sequences and large data sets. +Everything in `itertools` returns a **lazy iterator**, so values are produced on demand rather than built up in a list. That is what makes the module usable with infinite sequences and large data sets. Here are some of its most useful functions: ```python import itertools @@ -2146,11 +2145,11 @@ for key, group in itertools.groupby(data, key=lambda pair: pair[0]): print(list(itertools.accumulate([1, 2, 3, 4]))) # [1, 3, 6, 10] ``` -Note that `count` and `cycle` are **infinite** — never call `list()` on them directly, or the program will hang and eventually exhaust memory. Pair them with `islice`, `zip`, or an explicit `break`. Note too the common `groupby` trap: it only groups _adjacent_ items, so sort by the same key first if you want all matching items grouped together. +Note that `count` and `cycle` are **infinite** — never call `list()` on them directly, or the program will hang and eventually run out of memory. Pair them with `islice`, `zip`, or an explicit `break`. Also watch out for the common `groupby` trap: it only groups _adjacent_ items, so sort by the same key first if you want all matching items grouped together. -## 49- what does itertools.islice do? +## 49- What does `itertools.islice` do? -`itertools.islice` is a function that returns an iterator that returns selected elements from the input iterator. It works by slicing the input iterator and returning an iterator that produces the sliced elements. +`itertools.islice` works like slicing (`[start:stop:step]`), but for any iterable: it returns an iterator that produces only the selected elements of its input. Here's an example of how you can use `itertools.islice`: @@ -2182,7 +2181,7 @@ print(list(first_five_evens)) # [0, 2, 4, 6, 8] Two practical caveats: `islice` does not accept negative indices (it cannot count from the end without consuming everything), and it **consumes** the underlying iterator, so elements it skips are gone for good. -You can also use `itertools.islice` to slice at a specific starting and ending position with a specific step size. For example, `itertools.islice(numbers, 2, 6, 1)` returns an iterator producing the elements at indices `2` through `5`. Below are examples of output based on various inputs: +The signature mirrors slicing: `islice(iterable, stop)` or `islice(iterable, start, stop[, step])`. For example, `itertools.islice(numbers, 2, 6)` returns an iterator producing the elements at indices `2` through `5`. Here are more examples of the output for various inputs: ```python # itertools.islice(iterable, stop) @@ -2194,7 +2193,7 @@ You can also use `itertools.islice` to slice at a specific starting and ending p # islice('ABCDEFG', 0, None, 2) --> A C E G ``` -## 50- Why this code will never stop? +## 50- Why will this code never stop? ```python i = 0 @@ -2241,14 +2240,15 @@ while i != Decimal("1"): i += Decimal("0.1") # terminates: Decimal("0.1") is exact ``` -This is not a Python quirk — it is the IEEE 754 standard used by essentially every language. +This is not a Python quirk — it follows from the IEEE 754 floating-point standard, which essentially every programming language uses. ## 51- What is the output of this code, and why? ```python import datetime from time import sleep -def my_time(time_now = datetime.datetime.now()): + +def my_time(time_now=datetime.datetime.now()): return time_now print(my_time()) @@ -2260,12 +2260,13 @@ The output of this code will be two **identical** timestamps, even though three As a result, the value of `time_now` is fixed at the moment the function is defined. Both calls return the very same `datetime` object (`my_time() is my_time()` is `True`), which is why the printed values match to the microsecond. This is the classic "mutable/evaluated-once default argument" pitfall: default values live on the function object itself (inspect `my_time.__defaults__`), so anything computed or mutable there is shared across all calls. -To fix this issue, you could remove the default value for the `time_now` parameter, and set it to the current time inside the function using `datetime.datetime.now()`, like this: +To fix this, use `None` as the default and compute the current time inside the function, so it is evaluated on every call: ```python import datetime from time import sleep -def my_time(time_now = None): + +def my_time(time_now=None): if time_now is None: time_now = datetime.datetime.now() return time_now @@ -2275,11 +2276,11 @@ sleep(3) print(my_time()) ``` -## 52- Can we chain Multiple Decorators in Python? +## 52- Can we chain multiple decorators in Python? -Yes, you can chain multiple decorators in Python. Decorators are functions that are used to modify the behavior of another function. They are applied using the `@` symbol and can be used to add additional functionality to a function without modifying the function's source code. +Yes, you can chain multiple decorators in Python. A decorator wraps a function to add behavior without changing the function's source code (see question 26). -To chain multiple decorators, you can simply apply them one after the other, using the `@` symbol, like this: +To chain decorators, stack them one per line above the function definition, like this: ```python @decorator1 @@ -2295,7 +2296,7 @@ Stacked decorators are applied **bottom-up**: the decorator closest to the `def` function = decorator1(decorator2(decorator3(function))) ``` -So at definition time `decorator3` is applied first and `decorator1` last — `decorator1` ends up as the outermost wrapper. At **call** time the order reverses: the outermost wrapper (`decorator1`) runs first, then delegates inward. Keeping these two orders straight — bottom-up application, top-down execution — is the point interviewers usually probe. +So, at definition time, `decorator3` is applied first and `decorator1` last — `decorator1` ends up as the outermost wrapper. At **call** time the order reverses: the outermost wrapper (`decorator1`) runs first, then delegates inward. Keeping these two orders straight — bottom-up application, top-down execution — is what interviewers usually test. ```python def decorator1(func): @@ -2326,9 +2327,9 @@ function() Here `function = decorator1(decorator2(function))`: `decorator2` wrapped the function first, `decorator1` wrapped the result. When called, `decorator1`'s wrapper prints first, calls `decorator2`'s wrapper, which finally calls the original function — producing the top-down output shown. As with any decorator, apply `functools.wraps` in real code so the chain preserves the original function's name and docstring. -## 53- Build a recursive function using python +## 53- Build a recursive function using Python -To build a recursive function in Python, you will need to define a function that calls itself with a modified version of its input. Here's an example of how you can build a recursive function to calculate the factorial of a number: +To build a recursive function in Python, define a function that calls itself on a smaller version of its input. Here's an example of how you can build a recursive function to calculate the factorial of a number: ```python def factorial(n): @@ -2347,7 +2348,7 @@ In this example, the `factorial` function is defined to take a single argument ` Two details matter more than they look: - **The base case must cover every terminating input.** With the common `if n == 1` base case, `factorial(0)` never terminates — each call passes a smaller negative number until Python raises `RecursionError: maximum recursion depth exceeded`. `0! == 1` by definition, so the guard should be `n <= 1`, with an explicit `ValueError` for negatives. -- **Python does not optimise tail calls**, and the default recursion limit is about 1000 frames (`sys.getrecursionlimit()`), so `factorial(3000)` raises `RecursionError` even though the maths is fine. For depths like that, prefer an iterative version — or `math.factorial`, which is what production code should call anyway. +- **Python does not optimize tail calls** (it never reuses the current stack frame, even when a function ends by calling itself), and the default recursion limit is about 1000 frames (`sys.getrecursionlimit()`), so `factorial(3000)` raises `RecursionError` even though the math is fine. For depths like that, prefer an iterative version — or `math.factorial`, which is what production code should call anyway. ```python # 1st call: return n * factorial(4) "n = 5" @@ -2363,9 +2364,9 @@ Two details matter more than they look: # 1st call: 5 * 24 = 120 (5! -> 120) ``` -## 54- How to implement a binary search tree using Python? +## 54- How do you implement a binary search tree in Python? -To implement a binary search tree (BST) in Python, you will need to create a `Node` class to represent the nodes of the tree, and a `BST` class to represent the `BST` itself. +A binary search tree (BST) is a tree in which every node has at most two children, and the values are kept in order: everything in a node's left subtree is smaller than the node's value, and everything in its right subtree is greater than or equal to it. To implement one in Python, create a `Node` class to represent each node and a `BST` class to represent the tree itself. Here is an example of how you could implement a `Node` class in Python: @@ -2377,9 +2378,9 @@ class Node: self.right = None ``` -This `Node` class has three instance variables: `value`, `left`, and `right`. The `value` variable stores the value of the `node`, and the `left` and `right` variables are references to the `left` and `right` child nodes, respectively. +This `Node` class has three instance attributes: `value`, `left`, and `right`. `value` stores the node's value, and `left` and `right` hold references to the left and right child nodes (or `None` if there is no child). -To implement the `BST` itself, you will need to create a `BST` class that has methods for inserting and searching for nodes in the `tree`. Here is an example of how you could implement a `BST` class in Python: +Next, create a `BST` class with methods for inserting and searching for values in the tree. Here is an example: ```python class BST: @@ -2418,9 +2419,9 @@ class BST: return False ``` -The `BST` class has a `root` variable to store the `root` node of the `tree`, and two methods: `insert` and `search`. The `insert` method is used to insert a new node into the tree, and the `search` method is used to search for a node with a particular value. +The `BST` class has a `root` attribute that stores the root node, and two methods: `insert`, which adds a new node in the correct position, and `search`, which looks for a node with a particular value. Both walk down from the root, going left when the value is smaller than the current node's value and right otherwise. -To use these classes, you can create a new `BST` object and use the `insert` method to add nodes to the tree. You can then use the search method to search for a particular value in the tree. If a node with the specified `value` is found, the method returns `True`, otherwise it returns `False`. (As written, duplicates go to the right subtree; stating that policy explicitly is worth a sentence in an interview.) +To use these classes, create a `BST` object and call `insert` to add values to the tree. Then call `search` to look for a particular value: it returns `True` if a node with that value is found; otherwise, it returns `False`. (As written, duplicates go to the right subtree; stating that policy explicitly is worth a sentence in an interview.) An in-order traversal visits a BST's values in sorted order, which is both the standard way to read the tree back and a quick sanity check of the invariant: @@ -2440,9 +2441,11 @@ print(tree.search(6)) # True print(tree.search(7)) # False ``` -Complexity: `insert` and `search` are **O(h)** where `h` is the tree height — O(log n) on average for random insertion order, but **O(n) in the worst case**, because inserting already-sorted data degenerates the tree into a linked list. Self-balancing variants (AVL, red-black trees — the structure behind `sorted containers` in other languages) guarantee O(log n); Python's standard library instead offers the `bisect` module over a sorted list for the common cases. +Complexity: `insert` and `search` are **O(h)** where `h` is the tree height — O(log n) on average for random insertion order, but **O(n) in the worst case**, because inserting already-sorted data degenerates the tree into a linked list. Self-balancing variants (AVL and red-black trees — the structure behind sorted maps such as Java's `TreeMap` and C++'s `std::map`) guarantee O(log n). Python's standard library has no balanced tree; for most use cases, the `bisect` module on a sorted list (or the third-party `sortedcontainers` package) covers the same needs. + +## 55- How do you implement binary search in Python? -## 55- How to implement a binary search using Python? +Binary search finds an item in a **sorted** list by comparing the target with the middle element and discarding the half that cannot contain it, over and over, until the item is found or nothing is left. Here is a recursive implementation: ```python # Returns index of x in arr if present, else -1 @@ -2453,16 +2456,16 @@ def binary_search(arr, low, high, x): mid = (high + low) // 2 - # If element is present at the middle itself + # If the element is present at the middle itself if arr[mid] == x: return mid - # If element is smaller than mid, then it can only - # be present in left subarray + # If the element is smaller than arr[mid], it can only + # be present in the left half elif arr[mid] > x: return binary_search(arr, low, mid - 1, x) - # Else the element can only be present in right subarray + # Otherwise, it can only be present in the right half else: return binary_search(arr, mid + 1, high, x) @@ -2471,11 +2474,11 @@ def binary_search(arr, low, high, x): return -1 # Test array -arr = [ 2, 3, 4, 10, 40 ] +arr = [2, 3, 4, 10, 40] x = 10 # Function call -result = binary_search(arr, 0, len(arr)-1, x) +result = binary_search(arr, 0, len(arr) - 1, x) if result != -1: print("Element is present at index", str(result)) @@ -2483,13 +2486,13 @@ else: print("Element is not present in array") ``` -The non-negotiable precondition: **the input must already be sorted** — binary search on unsorted data silently returns wrong answers. Each comparison halves the search space, giving O(log n) time; this recursive version also uses O(log n) stack space, while a `while`-loop version brings that down to O(1). +The essential precondition: **the input must already be sorted** — binary search on unsorted data silently returns wrong answers. Each comparison halves the search space, giving O(log n) time; this recursive version also uses O(log n) stack space, while a `while`-loop version brings that down to O(1). -In production code, reach for the standard library instead of hand-rolling: `bisect.bisect_left(arr, x)` returns the insertion point in O(log n), and `arr[i] == x` at that index confirms membership. +In production code, use the standard library instead of writing your own: `bisect.bisect_left(arr, x)` returns the insertion point in O(log n), and `arr[i] == x` at that index confirms membership. -## 56- How to implement a Linked list using Python? +## 56- How do you implement a linked list in Python? -You will need to define a `Node` class to represent the nodes of the linked list, and a `LinkedList` class to represent the linked list itself. +A linked list is a chain of nodes, where each node holds a value and a reference to the next node. To implement one, define a `Node` class to represent each node and a `LinkedList` class to represent the list itself. ```python class Node: @@ -2498,9 +2501,9 @@ class Node: self.next = None ``` -This `Node` class has two instance variables: `value` and `next`. The `value` variable holds the value of the `node`, and the `next` variable holds a reference to the `next` node in the linked list. +This `Node` class has two instance attributes: `value` and `next`. `value` holds the node's value, and `next` holds a reference to the next node in the list (or `None` for the last node). -Next, you can define the `LinkedList` class, which will contain methods for inserting nodes into the linked list and searching for specific values: +Next, define the `LinkedList` class, with methods for appending values and traversing the list: ```python class LinkedList: @@ -2524,7 +2527,7 @@ class LinkedList: current_node = current_node.next ``` -The `LinkedList` class has two methods: `append` and `traverse`. The `append` method takes a value as an argument, creates a new `Node` object with that value, and appends it to the end of the linked list. The `traverse` method traverses the linked list and prints the value of each node to the console. +The `LinkedList` class has two methods: `append` and `traverse`. `append` creates a new `Node` with the given value and attaches it to the end of the list. `traverse` walks the list from the head and prints the value of each node. You can use these classes to create and manipulate a linked list like this: @@ -2536,11 +2539,11 @@ my_list.append(3) my_list.traverse() # Output: 1 2 3 (one value per line) ``` -Note that this `append` walks the whole list each time, making it O(n) per call; keeping a `tail` reference alongside `head` makes appends O(1). The general trade-off versus a Python `list`: linked lists give O(1) insertion/removal at a known node without shifting elements, but O(n) indexed access, worse cache behaviour, and per-node object overhead. That is why `collections.deque` (a doubly-linked structure of blocks, with O(1) appends and pops at both ends) is almost always the right practical choice, and hand-written linked lists appear mainly in interviews. +Note that this `append` walks the whole list each time, making it O(n) per call; keeping a `tail` reference alongside `head` makes appends O(1). The general trade-off versus a Python `list`: linked lists give O(1) insertion/removal at a known node without shifting elements, but O(n) indexed access, worse CPU-cache behavior (nodes are scattered in memory instead of stored side by side), and per-node object overhead. That is why `collections.deque` (a doubly linked list of fixed-size blocks, with O(1) appends and pops at both ends) is almost always the right practical choice, and hand-written linked lists appear mainly in interviews. -## 57- what is `collections.OrderedDict`? +## 57- What is `collections.OrderedDict`? -it is a class in the Python `collections` module that provides an ordered dictionary implementation. Like a regular dictionary, an `OrderedDict` stores key-value pairs, but it remembers the order which the keys were added. +`OrderedDict` is a dictionary subclass in the `collections` module that remembers the order in which keys were inserted. Like a regular dictionary, it stores key-value pairs. Here's an example of how you can use an `OrderedDict`: @@ -2567,7 +2570,7 @@ for key, value in d.items(): # d 4 ``` -The `OrderedDict` maintains the order in which the keys were added. +The items come out in the order in which they were inserted. The senior-level nuance: since **Python 3.7, plain `dict` also preserves insertion order** as a language guarantee (and did so as an implementation detail in CPython 3.6). So "remembers insertion order" is no longer a reason to reach for `OrderedDict`. What it still offers over `dict`: @@ -2580,14 +2583,14 @@ The senior-level nuance: since **Python 3.7, plain `dict` also preserves inserti dict(a=1, b=2) == dict(b=2, a=1) # True ``` -- **`move_to_end(key, last=True)`** — reposition a key at either end in O(1), the operation that makes an LRU cache trivial to build (see the LRU question later in this file). +- **`move_to_end(key, last=True)`** — reposition a key at either end in O(1), the operation that makes an LRU cache easy to build (see question 111). - **`popitem(last=False)`** — pop from _either_ end; `dict.popitem()` only pops the most recently inserted item. -Internally it maintains a doubly-linked list alongside the hash table to support those reordering operations, so it costs more memory than a plain dict. Today the practical rule is: use `dict` by default, and `OrderedDict` only when you need reordering operations or order-sensitive equality. +Internally, it maintains a doubly linked list alongside the hash table to support those reordering operations, so it costs more memory than a plain dict. Today the practical rule is: use `dict` by default, and `OrderedDict` only when you need reordering operations or order-sensitive equality. -## 58- what is `collections.defaultdict`? +## 58- What is `collections.defaultdict`? -it is a class in the Python `collections` module — a `dict` subclass that takes a **default factory**: a zero-argument callable invoked to produce a value whenever a missing key is accessed. Note the distinction: you pass a _callable that builds_ the default (`int`, `list`, or your own function), not a default value itself. `int()` returns `0`, `list()` returns `[]`, which is where the defaults come from. +`defaultdict` is a `dict` subclass in the `collections` module that takes a **default factory**: a zero-argument callable that is called to produce a value whenever a missing key is accessed. Note the distinction: you pass a _callable that builds_ the default (`int`, `list`, or your own function), not a default value itself. `int()` returns `0`, `list()` returns `[]`, which is where the defaults come from. Here's an example of how you can use a `defaultdict`: @@ -2609,7 +2612,7 @@ print(d['d']) # Output: 0 print(d) # Output: defaultdict(, {'a': 1, 'b': 2, 'c': 3, 'd': 0}) ``` -Notice the last line: accessing `d['d']` did not just _return_ `0`, it **inserted** the key `'d'` with value `0`. Reads of missing keys mutate a `defaultdict`, which can surprise you if you probe it with `in`-style logic afterwards. +Notice the last line: accessing `d['d']` did not just _return_ `0`, it **inserted** the key `'d'` with value `0`. Reading a missing key modifies a `defaultdict`, which can surprise you if you later check for that key with `in`. The factory is triggered **only by `d[key]` lookups** (the `__missing__` hook). `d.get('missing')` still returns `None`, and `'missing' in d` is still `False` — neither touches the factory. @@ -2630,17 +2633,17 @@ print(by_letter) # 'c': ['cherry']}) ``` -Any zero-argument callable works as the factory — including a `lambda` for a non-trivial default (`defaultdict(lambda: "N/A")`) or even a nested `defaultdict` for auto-vivifying trees. With a plain `dict`, the closest equivalents are `d.setdefault(key, []).append(...)` or `dict.get(key, default)`, both of which are noisier in a loop. +Any zero-argument callable works as the factory — including a `lambda` for a custom default (`defaultdict(lambda: "N/A")`) or even a nested `defaultdict` for trees whose levels are created automatically on first access. With a plain `dict`, the closest equivalents are `d.setdefault(key, []).append(...)` or `dict.get(key, default)`, both of which are more verbose in a loop. ## 59- Can we implement an `array` using Python? -Yes! By using the `array` module. Python’s `array` module provides space-efficient storage of basic C-style data types like **`bytes, 32-bit integers, floating-point numbers, and so on`**. +Yes, with the standard library's `array` module. It provides space-efficient storage of basic C-style data types, such as bytes, 32-bit integers, and floating-point numbers. -Arrays created with the `array.array` class are mutable and behave similarly to lists except for one important difference: they’re **`typed arrays`** constrained to a single data type. +Arrays created with the `array.array` class are mutable and behave much like lists, with one important difference: they are **typed arrays**, restricted to a single data type. -Because of this constraint, `array.array` objects with many elements are more space efficient than `lists` and `tuples`. The elements stored in them are tightly packed, and this can be useful if you need to store many elements of the same type. +Because of this restriction, `array.array` objects with many elements are more space-efficient than lists and tuples. The elements are stored as raw C values packed side by side, rather than as separate Python objects, which helps when you need to store many values of the same type. -Also, arrays support many of the same methods as regular lists, and you might be able to use them as a drop-in replacement without requiring other changes to your application code. +Arrays also support many of the same methods as regular lists, so you can often use them as a drop-in replacement without other changes to your code. For numerical work, NumPy arrays are the usual choice, since they add fast vectorized operations on top of compact storage. ```python >>> import array @@ -2674,15 +2677,15 @@ TypeError: must be real number, not str ## 60- What is the `bytes` type? -The `bytes` type is an immutable sequence of _bytes_. It is similar to the _str_ type, but it is meant to hold raw binary data rather than Unicode text. +The `bytes` type is an immutable sequence of _bytes_ (integers from 0 to 255). It is similar to the `str` type, but it holds raw binary data rather than Unicode text. -You can create a `bytes` object by prefixing a string with the b character and enclosing it in quotes, like this: +You can create a `bytes` object by writing a string literal with a `b` prefix, like this: ```python b = b'Hello, world!' ``` -You can also create a `bytes` object from a list of integers using the `bytes` function: +You can also create a `bytes` object from a list of integers using the `bytes()` constructor: ```python b = bytes([104, 101, 108, 108, 111]) # b'hello' @@ -2708,12 +2711,12 @@ Note the asymmetry that trips people up: **indexing a `bytes` object yields an ` Two related points complete the picture: -- **`bytes` vs `str` is the binary/text boundary.** You convert between them explicitly with an encoding: `"héllo".encode("utf-8")` produces `bytes`, and `data.decode("utf-8")` produces `str`. Mixing them (`b'a' + 'a'`) raises `TypeError` — Python 3 refuses to guess an encoding, which is precisely what made Python 2's implicit conversions a bug factory. +- **`bytes` vs `str` is the binary/text boundary.** You convert between them explicitly with an encoding: `"héllo".encode("utf-8")` produces `bytes`, and `data.decode("utf-8")` produces `str`. Mixing them (`b'a' + 'a'`) raises `TypeError` — Python 3 refuses to guess an encoding, because Python 2's implicit guessing was a constant source of bugs. - **`bytearray` is the mutable counterpart** — same interface, but you can modify it in place (`ba[0] = 72`), which matters when building or patching binary buffers without copying. -You meet `bytes` at every I/O edge: files opened in `'rb'` mode, sockets, `subprocess` pipes, HTTP bodies, and hashing (`hashlib` consumes bytes). +You work with `bytes` whenever data crosses an I/O boundary: files opened in `'rb'` mode, sockets, `subprocess` pipes, HTTP bodies, and hashing (`hashlib` consumes bytes). -## 61- How to concatenate tuples in python? +## 61- How do you concatenate tuples in Python? You can use the `+` operator. For example: @@ -2725,15 +2728,15 @@ tuple3 = tuple1 + tuple2 print(tuple3) # Output: (1, 2, 3, 4, 5, 6) ``` -This will create a new tuple that contains the elements of `tuple1` followed by the elements of `tuple2`. Unpacking gives the same result and generalises to any number of inputs, including other iterables: +This will create a new tuple that contains the elements of `tuple1` followed by the elements of `tuple2`. Unpacking gives the same result and generalizes to any number of inputs, including other iterables: ```python tuple3 = (*tuple1, *tuple2) # (1, 2, 3, 4, 5, 6) ``` -Keep in mind that tuples are **immutable**, which means that you cannot modify an existing `tuple`. Even `tuple1 += tuple2` does not mutate anything — it builds a brand-new tuple and rebinds the name, an O(n) copy each time. For that reason, concatenating many tuples in a loop is quadratic; collect into a `list` (or use `itertools.chain`) and convert once at the end instead. +Keep in mind that tuples are **immutable**, which means that you cannot modify an existing `tuple`. Even `tuple1 += tuple2` does not mutate anything — it builds a brand-new tuple and rebinds the name, an O(n) copy each time. For that reason, concatenating many tuples in a loop takes quadratic time; collect the items in a `list` (or use `itertools.chain`) and convert to a tuple once at the end instead. -## 62- How to join two `sets`? +## 62- How do you join two sets? To join two sets in Python, you can use the `union` method, which returns a new `set` that contains all the elements from both sets. @@ -2755,9 +2758,9 @@ set1.update(set2) print(set1) # Output: {1, 2, 3, 4, 5} ``` -The operator spellings are equally idiomatic: `set1 | set2` is union, and `set1 |= set2` updates in place. The same pairing covers the other set operations — `&`/`intersection`, `-`/`difference`, `^`/`symmetric_difference`. One practical difference worth knowing: the methods accept any iterable (`set1.union([4, 5])` works), while the operators require both operands to be sets. +The operator forms are equally idiomatic: `set1 | set2` is union, and `set1 |= set2` updates in place. The same pairing covers the other set operations — `&`/`intersection`, `-`/`difference`, `^`/`symmetric_difference`. One practical difference worth knowing: the methods accept any iterable (`set1.union([4, 5])` works), while the operators require both operands to be sets. -## 63- What is the difference between Python's list methods append and extend? +## 63- What is the difference between Python's list methods `append` and `extend`? The `append` method adds an element to the end of a list. It takes a single element as an argument and does not return a new list. @@ -2775,9 +2778,9 @@ list1.extend(list2) print(list1) # prints [1, 2, 3, 4, 5, 6, 7] ``` -Neither method creates a new list — both mutate the existing list in place and return `None`. The real performance distinction is simply how much work there is to do: `append` is amortised O(1) because it adds exactly one element, while `extend(iterable)` is O(k) for k elements added — and calling `extend` once is faster than calling `append` k times in a loop, since it avoids k method-call round-trips. +Neither method creates a new list — both mutate the existing list in place and return `None`. The real performance distinction is simply how much work there is to do: `append` is amortized O(1) because it adds exactly one element, while `extend(iterable)` is O(k) for k elements added — and calling `extend` once is faster than calling `append` k times in a loop, since it avoids k separate method calls. -The gotcha interviewers fish for is what happens when you pass a _list_ to `append`: +The gotcha interviewers look for is what happens when you pass a _list_ to `append`: ```python lst = [1, 2, 3] @@ -2791,26 +2794,28 @@ print(lst) # [1, 2, 3, 4, 5] <- the elements are added individually Also note that `extend` accepts any iterable — a tuple, set, generator, or string. That last one is a classic accident: `lst.extend("ab")` adds `'a'` and `'b'` as two separate elements. `lst += iterable` is equivalent to `extend` (in-place), whereas `lst + other` builds a new list and requires both operands to be lists. -## 64- How to implement bubble sort in Python? +## 64- How do you implement bubble sort in Python? + +Bubble sort repeatedly walks through the list and swaps neighboring elements that are in the wrong order, until a full pass makes no swaps. With each pass, the largest remaining value "bubbles up" to its final position at the end. ```python def bubble_sort(lst): - # Set swap to True to enter the loop - swap = True - # Repeat the loop until no swaps are needed - while swap: - # Set swap to False to start the loop - swap = False - # Iterate through the list - for i in range(len(lst) - 1): - # Check if the current element is greater than the next element - if lst[i] > lst[i + 1]: - # Swap the elements - lst[i], lst[i + 1] = lst[i + 1], lst[i] - # Set swap to True to continue the loop - swap = True - # Return the sorted list - return lst + # Set swap to True to enter the loop + swap = True + # Repeat the loop until no swaps are needed + while swap: + # Assume the list is sorted until a swap proves otherwise + swap = False + # Iterate through the list + for i in range(len(lst) - 1): + # Check if the current element is greater than the next element + if lst[i] > lst[i + 1]: + # Swap the elements + lst[i], lst[i + 1] = lst[i + 1], lst[i] + # Set swap to True to continue the loop + swap = True + # Return the sorted list + return lst ``` **Function call:** @@ -2826,66 +2831,68 @@ print(sorted_list) # [1, 2, 5, 8, 9] - Average case: **O(n2)** - Worst case: **O(n2)** -Bubble sort is a teaching algorithm; its one redeeming property is the O(n) early exit on nearly-sorted data. In real code, `sorted()` / `list.sort()` use Timsort — a hybrid stable sort that is O(n log n) worst case and also exploits existing order for O(n) best case. +Bubble sort is a teaching algorithm; its one redeeming property is the O(n) early exit on nearly sorted data. In real code, `sorted()` and `list.sort()` use Timsort — a hybrid, stable sort that is O(n log n) in the worst case and takes advantage of existing order for an O(n) best case. + +## 65- How do you implement heap sort in Python? -## 65- How to implement Heap sort in Python? +Heap sort first rearranges the list into a _max-heap_: a binary tree stored inside the list, in which every parent is greater than or equal to its children, so the largest element is always at the root (index 0). It then repeatedly removes the root and restores the heap property. The version below collects the removed elements in a new list: ```python def heap_sort(lst): - # Create an empty list to store the extracted elements - sorted_list = [] - # Convert the input list into a max heap - heapify(lst) - # Keep extracting the root element (maximum value) from the heap - # until it is empty - while lst: - # Extract the root element from the heap and append it to the - # sorted list - sorted_list.append(heappop(lst)) - # The elements were extracted largest-first, so reverse for ascending order - sorted_list.reverse() - # Return the sorted list - return sorted_list + # Create an empty list to store the extracted elements + sorted_list = [] + # Convert the input list into a max heap + heapify(lst) + # Keep extracting the root element (maximum value) from the heap + # until it is empty + while lst: + # Extract the root element from the heap and append it to the + # sorted list + sorted_list.append(heappop(lst)) + # The elements were extracted largest-first, so reverse for ascending order + sorted_list.reverse() + # Return the sorted list + return sorted_list def heapify(lst): - # Start from the last parent node - start = (len(lst) - 2) // 2 - # Sift down each node to create a max heap - while start >= 0: - sift_down(lst, start, len(lst) - 1) - start -= 1 + # Start from the last parent node + start = (len(lst) - 2) // 2 + # Sift down each node to create a max heap + while start >= 0: + sift_down(lst, start, len(lst) - 1) + start -= 1 def sift_down(lst, start, end): - # Set the root as the starting element - root = start - # While the root has a child - while root * 2 + 1 <= end: - # Set the child as the root's left child - child = root * 2 + 1 - # If the child has a sibling and the sibling is greater than the - # child, set the sibling as the child - if child + 1 <= end and lst[child] < lst[child + 1]: - child += 1 - # If the child is greater than the root, swap them - if lst[root] < lst[child]: - lst[root], lst[child] = lst[child], lst[root] - # Set the child as the new root - root = child - # If no swap is needed, exit the loop - else: - return + # Set the root as the starting element + root = start + # While the root has a child + while root * 2 + 1 <= end: + # Set the child as the root's left child + child = root * 2 + 1 + # If the child has a sibling and the sibling is greater than the + # child, set the sibling as the child + if child + 1 <= end and lst[child] < lst[child + 1]: + child += 1 + # If the child is greater than the root, swap them + if lst[root] < lst[child]: + lst[root], lst[child] = lst[child], lst[root] + # Set the child as the new root + root = child + # If no swap is needed, exit the loop + else: + return def heappop(lst): - # Save the root (maximum value) and the last element - root = lst[0] - last = lst.pop() - # If the heap is not empty, set the last element as the root and - # sift it down - if lst: - lst[0] = last - sift_down(lst, 0, len(lst) - 1) - # Return the root - return root + # Save the root (maximum value) and the last element + root = lst[0] + last = lst.pop() + # If the heap is not empty, set the last element as the root and + # sift it down + if lst: + lst[0] = last + sift_down(lst, 0, len(lst) - 1) + # Return the root + return root ``` **Function call:** @@ -2915,28 +2922,29 @@ print(heap_sort_stdlib([5, 2, 8, 1, 9])) # [1, 2, 5, 8, 9] - Average case: **O(n log(n))** - Worst case: **O(n log(n))** -Building the heap is O(n); each of the n extractions costs O(log n). Heap sort is not stable, but it is the classic answer when you need guaranteed O(n log n) with O(1) auxiliary space (in the in-place variant). +Building the heap is O(n); each of the n extractions costs O(log n). Heap sort is not stable (equal elements may change their relative order), but it is the classic answer when you need guaranteed O(n log n) time with O(1) extra space (in the in-place variant). -## 66- How to implement Insertion sort in Python? +## 66- How do you implement insertion sort in Python? + +Insertion sort builds a sorted section at the front of the list, one element at a time: each new element is shifted left past all larger elements until it reaches its correct position — the way most people sort a hand of playing cards. ```python def insertion_sort(lst): - # Iterate through the list, starting from the second element - for i in range(1, len(lst)): - # Save the current element - current = lst[i] - # Set the position (j) as the index of the previous element - j = i - 1 - # Keep moving the current element to the left as long as it is - # smaller than the elements to its left - while j >= 0 and current < lst[j]: - lst[j + 1] = lst[j] - j -= 1 - # When the correct position is found, insert the current element - lst[j + 1] = current - # Return the sorted list - return lst - + # Iterate through the list, starting from the second element + for i in range(1, len(lst)): + # Save the current element + current = lst[i] + # Set the position (j) as the index of the previous element + j = i - 1 + # Keep moving the current element to the left as long as it is + # smaller than the elements to its left + while j >= 0 and current < lst[j]: + lst[j + 1] = lst[j] + j -= 1 + # When the correct position is found, insert the current element + lst[j + 1] = current + # Return the sorted list + return lst ``` **Function call:** @@ -2952,47 +2960,50 @@ print(sorted_list) # [1, 2, 5, 8, 9] - Average case: **O(n2)** - Worst case: **O(n2)** -## 67- How to implement Merge sort in Python? +Insertion sort is stable and very fast on small or nearly sorted inputs, which is why Timsort uses it internally for short runs. + +## 67- How do you implement merge sort in Python? + +Merge sort is a divide-and-conquer algorithm: it splits the list in half, sorts each half recursively, and then merges the two sorted halves into a single sorted list. ```python def merge_sort(lst): - # If the input list is empty or has only one element, return it - if len(lst) <= 1: - return lst - # Split the list into two halves - mid = len(lst) // 2 - left = lst[:mid] - right = lst[mid:] - # Recursively sort the two halves - left = merge_sort(left) - right = merge_sort(right) - # Merge the sorted halves and return the result - return merge(left, right) + # If the input list is empty or has only one element, return it + if len(lst) <= 1: + return lst + # Split the list into two halves + mid = len(lst) // 2 + left = lst[:mid] + right = lst[mid:] + # Recursively sort the two halves + left = merge_sort(left) + right = merge_sort(right) + # Merge the sorted halves and return the result + return merge(left, right) def merge(left, right): - # Create an empty list to store the merged elements - merged = [] - # Set the indices for the left and right lists - left_index = 0 - right_index = 0 - # While there are elements in both lists - while left_index < len(left) and right_index < len(right): - # If the left element is smaller or equal, take it first: taking from - # the LEFT on ties is what makes the sort stable - if left[left_index] <= right[right_index]: - merged.append(left[left_index]) - left_index += 1 - # If the right element is smaller, add it to the merged list - # and increment the right index - else: - merged.append(right[right_index]) - right_index += 1 - # Add the remaining elements (if any) to the merged list - merged.extend(left[left_index:]) - merged.extend(right[right_index:]) - # Return the merged list - return merged - + # Create an empty list to store the merged elements + merged = [] + # Set the indices for the left and right lists + left_index = 0 + right_index = 0 + # While there are elements in both lists + while left_index < len(left) and right_index < len(right): + # If the left element is smaller or equal, take it first: taking from + # the LEFT on ties is what makes the sort stable + if left[left_index] <= right[right_index]: + merged.append(left[left_index]) + left_index += 1 + # If the right element is smaller, add it to the merged list + # and increment the right index + else: + merged.append(right[right_index]) + right_index += 1 + # Add the remaining elements (if any) to the merged list + merged.extend(left[left_index:]) + merged.extend(right[right_index:]) + # Return the merged list + return merged ``` **Function call:** @@ -3008,22 +3019,24 @@ print(sorted_list) # [1, 2, 5, 8, 9] - Average case: **O(n log(n))** - Worst case: **O(n log(n))** -Merge sort is **stable** (equal elements keep their original relative order — guaranteed by the `<=` in `merge`) at the cost of O(n) auxiliary space. Python's built-in Timsort is a heavily optimised merge-sort/insertion-sort hybrid, which is why stability is guaranteed for `sorted()` and `list.sort()`. +Merge sort is **stable** (equal elements keep their original relative order — guaranteed by the `<=` in `merge`) at the cost of O(n) extra space. Python's built-in Timsort is a heavily optimized merge-sort/insertion-sort hybrid, which is why stability is guaranteed for `sorted()` and `list.sort()`. + +## 68- How do you implement quicksort in Python? -## 68- How to implement Quick Sort in Python? +Quicksort picks one element as the _pivot_, splits the remaining elements into those less than or equal to the pivot and those greater than it, and then sorts each group recursively. ```python def quick_sort(lst): - # If the input list has fewer than 2 elements, return it - if len(lst) < 2: - return lst - # Set the pivot as the first element in the list - pivot = lst[0] - # Create the lists for elements less than and greater than the pivot - less_than = [element for element in lst[1:] if element <= pivot] - greater_than = [element for element in lst[1:] if element > pivot] - # Recursively sort the two lists and return the result - return quick_sort(less_than) + [pivot] + quick_sort(greater_than) + # If the input list has fewer than 2 elements, return it + if len(lst) < 2: + return lst + # Set the pivot as the first element in the list + pivot = lst[0] + # Create the lists for elements less than and greater than the pivot + less_than = [element for element in lst[1:] if element <= pivot] + greater_than = [element for element in lst[1:] if element > pivot] + # Recursively sort the two lists and return the result + return quick_sort(less_than) + [pivot] + quick_sort(greater_than) ``` **Function call:** @@ -3039,25 +3052,27 @@ print(sorted_list) # [1, 2, 5, 8, 9] - Average case: **O(n log(n))** - Worst case: **O(n2)** -Two things to say about this elegant version: it is **not in place** (each level builds new lists, so O(n) extra space per level), and choosing the **first element as pivot** makes already-sorted input the worst case — every partition is maximally lopsided, degrading to O(n²) and deep recursion. Picking a random pivot (or median-of-three) makes that pathological case vanishingly unlikely; the in-place Lomuto/Hoare partition schemes are the standard follow-up whiteboard exercise. +Two things to say about this elegant version: it is **not in place** (each level builds new lists, so it needs O(n) extra space per level), and choosing the **first element as the pivot** makes already-sorted input the worst case — every split is as unbalanced as possible, degrading to O(n²) time and very deep recursion. Picking a random pivot (or the median of the first, middle, and last elements) makes that worst case extremely unlikely. The in-place Lomuto and Hoare partition schemes are the usual follow-up interview exercise. -## 69- How to implement Selection sort in Python? +## 69- How do you implement selection sort in Python? + +Selection sort repeatedly finds the smallest element in the unsorted part of the list and swaps it into place at the front of that part. ```python def selection_sort(lst): - # Iterate through the list, starting from the first element - for i in range(len(lst)): - # Set the minimum element as the current element - minimum = i - # Find the minimum element in the remaining list - for j in range(i + 1, len(lst)): - if lst[j] < lst[minimum]: - minimum = j - # If the minimum element is not the current element, swap them - if minimum != i: - lst[i], lst[minimum] = lst[minimum], lst[i] - # Return the sorted list - return lst + # Iterate through the list, starting from the first element + for i in range(len(lst)): + # Set the minimum element as the current element + minimum = i + # Find the minimum element in the remaining list + for j in range(i + 1, len(lst)): + if lst[j] < lst[minimum]: + minimum = j + # If the minimum element is not the current element, swap them + if minimum != i: + lst[i], lst[minimum] = lst[minimum], lst[i] + # Return the sorted list + return lst ``` **Function call:** @@ -3073,21 +3088,25 @@ print(sorted_list) # [1, 2, 5, 8, 9] - Average case: **O(n2)** - Worst case: **O(n2)** -## 70- How to implement Shell sort in Python? +Selection sort always makes O(n²) comparisons, even on sorted input, but it performs at most n − 1 swaps, which only matters when writing to memory is expensive. It is not stable in this form. + +## 70- How do you implement Shell sort in Python? + +Shell sort is a generalization of insertion sort. It first sorts elements that are far apart (a fixed _gap_ apart), then repeats with smaller and smaller gaps until the gap is 1, at which point it is a plain insertion sort over an almost-sorted list. ```python def shell_sort(arr): - gap = len(arr) // 2 - while gap > 0: - for i in range(gap, len(arr)): - temp = arr[i] - j = i - while j >= gap and arr[j - gap] > temp: - arr[j] = arr[j - gap] - j -= gap - arr[j] = temp - gap //= 2 - return arr + gap = len(arr) // 2 + while gap > 0: + for i in range(gap, len(arr)): + temp = arr[i] + j = i + while j >= gap and arr[j - gap] > temp: + arr[j] = arr[j - gap] + j -= gap + arr[j] = temp + gap //= 2 + return arr ``` **Function call:** @@ -3103,13 +3122,13 @@ print(sorted_list) # [1, 2, 3, 4, 5, 6] - Average case: depends on the gap sequence; roughly **O(n3/2)** for this one - Worst case: **O(n2)** -The shell sort is insertion sort performed over progressively smaller gaps, so far-apart elements move long distances early. Its complexity is governed entirely by the gap sequence — better sequences (Knuth's `3k+1`, Ciura's empirical sequence) improve the worst case to below O(n²) — and no gap sequence makes it beat O(n log n) sorts asymptotically. It is unstable, in place, and mostly of historical/embedded interest. +Because far-apart elements move long distances early, the final insertion-sort pass has little work left to do. Shell sort's complexity depends entirely on the gap sequence — better sequences (Knuth's `3k+1`, Ciura's empirical sequence) improve the worst case to below O(n²) — but no gap sequence makes it asymptotically faster than the O(n log n) sorts. It is unstable, sorts in place, and today is mostly of historical interest or used on memory-constrained embedded systems. -## 71- What are the commands that are used to copy an object in Python? +## 71- How do you copy an object in Python? -There are several ways to copy an object in Python. Here are some of the most common methods: +There are several ways to copy an object in Python. Here are the most common ones: -- Using the `copy` module: +- Using `copy.copy()`: ```python import copy @@ -3118,7 +3137,7 @@ There are several ways to copy an object in Python. Here are some of the most co This creates a shallow copy of the object. If the object contains references to other objects, the copy will contain references to the same objects as the original. -- Using the `deepcopy` function: +- Using `copy.deepcopy()`: ```python import copy @@ -3133,9 +3152,9 @@ There are several ways to copy an object in Python. Here are some of the most co new_object = old_object.copy() ``` - This creates a shallow copy of the object. This method is available for objects that support the `copy` protocol (e.g., lists, dictionaries, sets, etc.). + This creates a shallow copy of the object. Many built-in containers provide this method (e.g., lists, dictionaries, and sets). -A few equivalent idioms you will meet in real code: `lst[:]` and `list(lst)` shallow-copy a list, `dict(d)` and `d | {}` shallow-copy a dict, and since Python 3.3 sequences also expose `.copy()` directly. All of these are shallow — for anything nested, only `copy.deepcopy` duplicates the inner objects. Custom classes can hook into the mechanism by defining `__copy__` and `__deepcopy__`. +A few equivalent idioms you will meet in real code: `lst[:]` and `list(lst)` shallow-copy a list, `dict(d)` and `d | {}` shallow-copy a dict, and since Python 3.3, lists also have a `.copy()` method. All of these are shallow — for anything nested, only `copy.deepcopy` duplicates the inner objects. Custom classes can hook into the mechanism by defining `__copy__` and `__deepcopy__`. ## 72- What is the difference between deep and shallow copy? @@ -3171,17 +3190,17 @@ Practical notes a senior engineer is expected to add: - **Immutables are not really copied**: for ints, strings, and tuples of immutables, both copy forms may return the same objects — that is safe precisely because those objects can never change. - The multiplication trap is the same bug in disguise: `[[0] * 3] * 2` builds two references to _one_ inner list; use a comprehension (`[[0] * 3 for _ in range(2)]`) to get independent rows. -## 73- How can the ternary operators be used in Python? +## 73- How can the ternary operator be used in Python? -In Python, the ternary operator is known as the conditional operator or ternary conditional operator. It is an operator that takes three arguments: a condition, a result for the condition being true, and a result for the condition being false. +Python's ternary operator is officially called a _conditional expression_. It takes three operands: a condition, a value to use when the condition is true, and a value to use when it is false. -The syntax for the ternary operator is: +The syntax is: ```python result = expression1 if condition else expression2 ``` -Here, `expression1` and `expression2` are the results that are returned if the condition is true or false, respectively. +Here, `expression1` is the result if `condition` is true, and `expression2` is the result if it is false. Only the chosen expression is evaluated. Here's an example of how you can use the ternary operator to assign a value to a variable based on a condition: @@ -3204,7 +3223,9 @@ max_value = get_max_value(10, 20) print(max_value) ``` -In this example, the function `get_max_value()` returns `x` if `x` is greater than `y`, and returns `y` if `x` is not greater than `y`. When called with the arguments `(10, 20)`, the function will return `20`. +In this example, `get_max_value()` returns `x` if `x` is greater than `y`, and `y` otherwise. When called with the arguments `(10, 20)`, it returns `20`. + +Keep conditional expressions short. Nesting them (`a if x else b if y else c`) quickly hurts readability; use a regular `if` statement instead. ## 74- What will be the output of the code below? @@ -3214,7 +3235,7 @@ def extendList(val, list=[]): return list list1 = extendList(10) -list2 = extendList(123,[]) +list2 = extendList(123, []) list3 = extendList('a') print("list1 = %s" % list1) @@ -3224,20 +3245,21 @@ print("list3 = %s" % list3) **The output:** -```python +```text list1 = [10, 'a'] list2 = [123] list3 = [10, 'a'] ``` -- In the first call to `extendList()`, the default value of the list is used, which is an empty list `[]`. The value `10` is appended to this list, and the modified list is returned. This list is assigned to `list1`. +- In the first call to `extendList()`, the default list is used, which starts out empty. The value `10` is appended to it, and the list is returned and assigned to `list1`. + +- In the second call, a new empty list `[]` is passed explicitly, so the default is not used. The value `123` is appended to it, and that list is returned and assigned to `list2`. -- In the second call to `extendList()`, a new list `[123]` is passed as the value for the list parameter, so the default value is not used. The value `123` is appended to this list, and the modified list is returned and assigned to `list2`. +- In the third call, the default list is used again. But it is the _same_ list object that the first call modified, so it already contains `10`. The value `'a'` is appended to it, and the list is returned and assigned to `list3`. -- In the third call to `extendList()`, the default value of the list is used again. This time, the default value is the list that was modified in the first call to the function, which contains the value `10`. The value `'a'` is appended to this list, and the modified list is returned and assigned to `list3`. - This behavior occurs because default values are evaluated when the function is defined, not when it is called. In this case, the default value of the list parameter is an empty list `[]`, which is evaluated when the `extendList()` function is defined. This means that the same `list` object is used as the default value for the `list` parameter every time the `extendList()` function is called, unless a different value is provided for the list parameter in the function call. +This happens because default values are evaluated only once, when the `def` statement runs, not each time the function is called. The same list object is therefore reused as the default in every call that doesn't pass its own list. `list1` and `list3` are actually the _same object_ (`list1 is list3` is `True`), which is why both print `[10, 'a']`, even though `list1` was printed after the third call. -The definition of the `extendList` function could be modified as follows, though, to always begin a new list when no `list` argument is specified, which is more likely to have been the desired behavior: +To start a new list whenever no `list` argument is given, which is almost certainly the intended behavior, change the definition as follows: ```python def extendList(val, list=None): @@ -3253,16 +3275,16 @@ The `list=None` / `if list is None` pattern is the standard idiom for mutable de ```python def multipliers(): - return [lambda x : i * x for i in range(4)] + return [lambda x: i * x for i in range(4)] print([m(2) for m in multipliers()]) ``` The output of the above code will be `[6, 6, 6, 6]`. -The reason for this is that Python’s closures are late binding. This means that the values of variables used in closures are looked up at the time the inner function is called. So as a result, when any of the functions returned by `multipliers()` are called, the value of `i` is looked up in the surrounding scope at that time. By then, regardless of which of the returned functions is called, the `for` loop has been completed, and `i` is left with its final value of 3. Therefore, every returned function multiplies the value it is passed by `3`, so since a value of `2` is passed in the above code, they all return a value of `6` (i.e., 3 x 2). +The reason is that Python's closures are _late-binding_: a closure looks up the values of the outer variables it uses when it is _called_, not when it is created. All four lambdas refer to the same variable `i`. By the time any of them is called, the loop has finished, and `i` holds its final value, `3`. So every lambda multiplies its argument by `3`, and since `2` is passed in, they all return `6` (3 × 2). -The standard fix exploits the fact that **default arguments are evaluated at definition time** (the very behaviour that causes the previous question's bug is the cure here) — bind the current `i` as a default: +The standard fix relies on the fact that **default arguments are evaluated at definition time** (the very behavior that causes the previous question's bug is the cure here) — bind the current `i` as a default: ```python def multipliers(): @@ -3300,19 +3322,19 @@ print(Parent.x, Child1.x, Child2.x) 3 2 3 ``` -In Python, class variables are internally handled as dictionaries. If a variable name is not found in the dictionary of the current class, the class hierarchy (i.e., its parent classes) is searched until the referenced variable name is found (if the referenced variable name is not found in the class itself or anywhere in its hierarchy, an `AttributeError` occurs). +In Python, each class stores its attributes in its own namespace dictionary (`__dict__`). If a name is not found in the class's own dictionary, Python searches its parent classes, in method resolution order, until the name is found. If the name is not found anywhere in the hierarchy, an `AttributeError` is raised. -Therefore, setting `x = 1` in the `Parent` class makes the class variable `x` (with a value of 1) referenceable in that class and any of its children. That’s why the first `print` statement outputs `1 1 1`. +Therefore, setting `x = 1` in the `Parent` class makes `x` visible in `Parent` and in all of its children, which don't have an `x` of their own. That's why the first `print` statement outputs `1 1 1`. -Subsequently, if any of its child classes overrides that value (for example, when we execute the statement `Child1.x = 2`), then the value is changed in that child only. That’s why the second `print` statement outputs `1 2 1`. +Next, `Child1.x = 2` creates a new attribute in `Child1`'s own dictionary, which shadows the parent's `x` for `Child1` only. `Parent` and `Child2` are unaffected. That's why the second `print` statement outputs `1 2 1`. -Finally, if the value is then changed in the `Parent` (for example, when we execute the statement `Parent.x = 3`), that change is reflected also by any children that have not yet overridden the value (which in this case would be `Child2`). That’s why the third print statement outputs `3 2 3` +Finally, `Parent.x = 3` changes the value in `Parent`. The change is also seen by every child that has not defined its own `x` (here, `Child2`). That's why the third `print` statement outputs `3 2 3`. -## 77- What is `__slots__` in python? +## 77- What is `__slots__` in Python? -It is a feature of Python classes that allows you to specify the attributes that an instance of the class should have. By default, Python classes create a dictionary for each instance to store its attributes. This dictionary takes up more memory than is necessary for most objects and can cause performance problems, especially for objects with a large number of attributes. +`__slots__` is a class attribute that declares, up front, the fixed set of attributes that instances of the class can have. By default, Python gives every instance its own dictionary (`__dict__`) to store its attributes. A dictionary is flexible but relatively memory-hungry, and that cost adds up when you create a large number of instances. -Using `__slots__` can help mitigate this issue by allowing you to specify exactly which attributes an instance should have, and the Python interpreter will use a more efficient representation for the instance. This can save memory and improve performance. +With `__slots__`, Python instead reserves a fixed slot for each listed attribute directly inside the instance and skips the per-instance dictionary. This saves memory and makes attribute access slightly faster. Here's an example of how you might use `__slots__` in a Python class: @@ -3325,7 +3347,7 @@ class Point: self.y = y ``` -In this example, the Point class has two attributes: `x` and `y`. By specifying these attributes in `__slots__`, we are telling the Python interpreter that instances of the Point class should only have these two attributes and no others. +In this example, the `Point` class has two attributes: `x` and `y`. Listing them in `__slots__` tells Python that instances of `Point` can have only these two attributes and no others. There are a few things to note about using `__slots__`: @@ -3333,13 +3355,13 @@ There are a few things to note about using `__slots__`: - If you define `__slots__` in a class, its instances will not have a `__dict__`, so assigning any attribute not listed raises `AttributeError` — and instances also lose `__weakref__` unless you include it in the slots. Under the hood, slotted attributes are implemented as descriptors backed by fixed storage in the instance, which is where both the memory saving and the small attribute-access speedup come from. -- **Inheritance is the classic trap**: if any class in the hierarchy lacks `__slots__` (including the case where a subclass simply doesn't declare one), its instances get a `__dict__` anyway and the memory benefit silently evaporates. Every class in the chain must declare `__slots__` (an empty tuple `__slots__ = ()` is fine for mixins), and each class should list only its _new_ attributes. +- **Inheritance is the classic trap**: if any class in the hierarchy lacks `__slots__` (including the case where a subclass simply doesn't declare one), its instances get a `__dict__` anyway, and the memory benefit silently disappears. Every class in the chain must declare `__slots__` (an empty tuple `__slots__ = ()` is fine for mixins), and each class should list only its _new_ attributes. -- Using `__slots__` means committing to a fixed attribute set, which trades away Python's usual dynamism — no monkey-patching attributes onto instances, and tools that expect `vars(obj)` will not work. Reserve it for classes instantiated in large numbers (think millions of points, nodes, or rows), where the per-instance saving is real; measure with `sys.getsizeof` + `tracemalloc` rather than assuming. For plain data records, `dataclasses.dataclass(slots=True)` (Python 3.10+) generates the slots for you, and `NamedTuple` is a slot-like immutable alternative. +- Using `__slots__` means committing to a fixed set of attributes, which gives up some of Python's usual flexibility — you can no longer add arbitrary attributes to instances at runtime, and tools that rely on `vars(obj)` will not work. Reserve it for classes instantiated in large numbers (think millions of points, nodes, or rows), where the per-instance saving is real; measure with `sys.getsizeof` and `tracemalloc` rather than assuming. For plain data records, `dataclasses.dataclass(slots=True)` (Python 3.10+) generates the slots for you, and `NamedTuple` is a slot-like immutable alternative. -## 78- What is `__contains__` in python? +## 78- What is `__contains__` in Python? -The `__contains__` method is a special method in Python that is used to implement the `in` operator. If a class defines a `__contains__` method, you can use the in operator to check if an instance of the class contains a particular value. +The `__contains__` method is a special method that implements the `in` operator. If a class defines `__contains__`, you can use `in` to check whether an instance of the class contains a particular value. Here's an example of how you might use the `__contains__` method in a Python class: @@ -3359,9 +3381,9 @@ print(255 in color) # prints True print(0 in color) # prints True - g and b are both 0 ``` -In this example, the `Color` class has a `__contains__` method that checks if the given value is one of the `r`, `g`, or `b` values of the `Color` instance. When you use the in operator with an instance of the `Color` class, it will call the `__contains__` method to determine whether the value is contained within the instance. +In this example, the `Color` class has a `__contains__` method that checks whether the given value is one of the instance's `r`, `g`, or `b` values. When you use `in` with a `Color` instance, Python calls `__contains__` to get the answer. -Worth knowing for the follow-up question: `in` works even **without** `__contains__`. If the method is absent, Python falls back to iterating the object (`__iter__`), comparing each element; failing that, it tries the old `__getitem__` integer-indexing protocol. Defining `__contains__` is therefore an optimisation and a semantic statement — e.g. `dict` and `set` implement it as an O(1) hash lookup rather than a linear scan. The result of `__contains__` is also interpreted as a boolean, and `not in` simply negates it. +Worth knowing for the follow-up question: `in` works even **without** `__contains__`. If the method is absent, Python falls back to iterating the object (`__iter__`), comparing each element; failing that, it tries the old `__getitem__` integer-indexing protocol. Defining `__contains__` is therefore both an optimization and a way to define what membership _means_ for your type — e.g., `dict` and `set` implement it as an O(1) hash lookup rather than a linear scan. The result of `__contains__` is also interpreted as a boolean, and `not in` simply negates it. ## 79- What is a "callable"? @@ -3434,27 +3456,29 @@ x ^= 10 # augmented assignment works too **Logical XOR** ("exactly one of the two is truthy") has no dedicated operator; the idioms are: -1. `bool(a) != bool(b)` — the clearest and most common -2. `bool(a) ^ bool(b)` — works because `bool` is an `int` subclass (`True ^ False` → `True`) -3. `(a and not b) or (not a and b)` — spelled out with logic operators +1. `bool(a) != bool(b)` — the clearest and most common. +2. `bool(a) ^ bool(b)` — works because `bool` is a subclass of `int` (`True ^ False` → `True`). +3. `(a and not b) or (not a and b)` — spelled out with logical operators. `operator.xor(a, b)` is simply a function version of `^`, so it is _bitwise_ — `operator.xor(2, 4)` is `6`, not a truth test. Wrap the arguments in `bool()` if you want it to behave logically. -## 81- What is introspection/reflection and does Python support it? +## 81- What is introspection/reflection, and does Python support it? -Introspection is the ability to examine an object at runtime. Python has a `dir()` function that supports examining the attributes of an object, `type()` to check the object type, `isinstance()`, etc. While introspection is a passive examination of the objects, reflection is a more powerful tool where we can modify objects at runtime and access them dynamically. E.g. +Yes, fully. Introspection is the ability of a program to examine objects at runtime: what type they are, which attributes and methods they have, and so on. In Python, `dir()` lists an object's attributes, `type()` returns its type, and `isinstance()` checks whether it is an instance of a given class. -- `setattr()` adds or modifies an object's attribute; -- `getattr()` gets the value of an attribute of an object. +Introspection only _examines_ objects. Reflection goes a step further: the program can also access and modify objects dynamically at runtime, using names that are only known while the program runs. For example: -It can even invoke functions dynamically - `getattr(my_obj, "my_func_name")()` +- `getattr()` reads an attribute by name (given as a string). +- `setattr()` adds or changes an attribute by name. -Rounding out the toolbox: `hasattr()` probes for an attribute, `vars(obj)` returns the instance `__dict__`, `id()` gives the object's identity, and the `inspect` module goes deeper — `inspect.signature()` reads a function's parameters, `inspect.getsource()` retrieves its source code, and `inspect.getmembers()` enumerates attributes with filters. This runtime openness is what makes frameworks possible: ORMs discover model fields, pytest finds `test_*` functions, and serialisers walk objects — all via introspection rather than code generation. +You can even call a method by name: `getattr(my_obj, "my_func_name")()`. + +Other useful tools: `hasattr()` checks whether an attribute exists, `vars(obj)` returns the instance `__dict__`, `id()` gives the object's identity, and the `inspect` module goes deeper — `inspect.signature()` reads a function's parameters, `inspect.getsource()` retrieves its source code, and `inspect.getmembers()` lists attributes, optionally filtered by kind. This runtime openness is what makes frameworks possible: ORMs discover model fields, pytest finds `test_*` functions, and serializers walk objects — all via introspection rather than code generation. ## 82- What will be the output of lines 2, 4, 6, and 8 from the following code, and why? ```python -list = [ [ ] ] * 5 +list = [[]] * 5 list # output? list[0].append(10) list # output? @@ -3467,7 +3491,7 @@ list # output? **The output:** ```python -list = [ [ ] ] * 5 +list = [[]] * 5 list # output: [[], [], [], [], []] list[0].append(10) list # output: [[10], [10], [10], [10], [10]] @@ -3477,7 +3501,7 @@ list.append(30) list # output: [[10, 20], [10, 20], [10, 20], [10, 20], [10, 20], 30] ``` -In the first line, `list` is initialized with 5 references to the **same** single empty list — `*` copies references, it does not clone objects. When you append `10` through the first reference, you are modifying the one shared inner list, which is why the change shows through all five elements. In the fourth line, you are appending to the _outer_ list, which is why `30` lands at the end rather than inside any inner list. +In the first line, `list` is initialized with 5 references to the **same** single empty list — `*` copies references, it does not clone objects. When you append `10` through the first reference, you are modifying the one shared inner list, which is why the change shows up in all five elements. Appending `20` through the second reference changes that same shared list again. On line 7, however, you are appending to the _outer_ list, which is why `30` lands at the end rather than inside any inner list. The fix is a comprehension, which evaluates `[]` freshly on every iteration: @@ -3487,11 +3511,11 @@ grid[0].append(10) print(grid) # [[10], [], [], [], []] ``` -The rule of thumb: `[x] * n` is fine when `x` is immutable (numbers, strings), and a latent bug when `x` is mutable. (The snippet also shadows the built-in `list` type by using it as a variable name — after line 1, `list()` no longer constructs lists in that scope. Avoid that in real code.) +The rule of thumb: `[x] * n` is fine when `x` is immutable (numbers, strings), and a hidden bug when `x` is mutable. (The snippet also shadows the built-in `list` type by using it as a variable name — after line 1, `list()` no longer constructs lists in that scope. Avoid that in real code.) -## 83- Write a function that prints the least integer that is not present in a given list and cannot be represented by the summation of the sub-elements of the list +## 83- Write a function that finds the smallest positive integer that cannot be represented as the sum of any subset of a given list -This is the classic "smallest unrepresentable sum" problem: given a list of positive integers, find the smallest positive integer that cannot be written as the sum of any subset of the list. There is a beautiful greedy O(n log n) solution: +This is the classic "smallest unrepresentable sum" problem: given a list of positive integers, find the smallest positive integer that cannot be written as the sum of any subset of the list. There is an elegant greedy solution that runs in O(n log n) time (the sort dominates; the scan itself is O(n)): ```python def find_least_integer(lst): @@ -3511,9 +3535,9 @@ def find_least_integer(lst): return reach ``` -The key insight is the invariant: after processing some prefix of the sorted list, if every integer in `[1, reach)` is achievable, then a new number `num <= reach` extends that range to `[1, reach + num)` — take any existing sum and optionally add `num`. But if `num > reach`, then `reach` itself can never be formed: all remaining numbers are at least `num`, which is already too big, so `reach` is the answer. +The key insight is the invariant: after processing the first few numbers of the sorted list, if every integer in `[1, reach)` is achievable, then a new number `num <= reach` extends that range to `[1, reach + num)` — take any existing sum and optionally add `num`. But if `num > reach`, then `reach` itself can never be formed: all remaining numbers are at least `num`, which is already too big, so `reach` is the answer. -Note that a simpler "walk until a number is missing" loop (incrementing only when `num == least_int`) answers a _different_ question — the smallest **missing** integer — and gets this problem wrong: for `[1, 1, 3]` it would return `2`, yet `1 + 1 = 2` is clearly representable; the true answer is `6` (`1+1+3 = 5`, and `6` cannot be formed). +Note that a simpler loop that counts upward until it finds a number missing from the list answers a _different_ question — the smallest **missing** integer — and gets this problem wrong: for `[1, 1, 3]`, it would return `2`, yet `1 + 1 = 2` is clearly representable. The true answer is `6`: the subset sums cover 1 through 5, and 6 cannot be formed. Here is an example of how to use the function: @@ -3548,7 +3572,7 @@ Here are three ways to reverse a list in Python: ```python lst = [1, 2, 3, 4, 5] reversed_lst = [] - for i in range(len(lst)-1, -1, -1): + for i in range(len(lst) - 1, -1, -1): reversed_lst.append(lst[i]) print(reversed_lst) # Output: [5, 4, 3, 2, 1] ``` @@ -3562,41 +3586,39 @@ lst = [1, 2, 3, 4, 5] for x in reversed(lst): # lazy iterator: no copy made at all print(x) -rev = list(reversed(lst)) # materialise when you actually need a list +rev = list(reversed(lst)) # materialize when you actually need a list ``` `reversed()` returns a lazy iterator over the existing list, so it costs O(1) memory — the right choice when you only need to _iterate_ backwards. Use `lst.reverse()` when you want the list itself permanently flipped, and `lst[::-1]` when you need a reversed copy as an expression (it also works on strings and tuples, which have no `.reverse()` method). ## 85- How does Python execute code? -When you run a Python program, the interpreter executes the code you have written in a sequence of steps. - -Here is a general outline of how the Python interpreter executes code: +When you run a Python program, CPython (the standard interpreter) processes your code in several steps: -1. The source is **parsed** — tokenised and built into an abstract syntax tree (AST); syntax errors surface at this stage, before anything runs. +1. The source is **parsed**: it is split into tokens and built into an abstract syntax tree (AST), a tree that represents the structure of the code. Syntax errors are reported at this stage, before anything runs. 2. The AST is **compiled to bytecode**, a compact instruction set for CPython's stack-based virtual machine (you can inspect it with the `dis` module). For imported modules the bytecode is cached on disk in `__pycache__` to skip recompilation next time. 3. The **Python virtual machine** (the eval loop) then executes the bytecode instruction by instruction. -4. As the interpreter executes the instructions, it may encounter statements that define variables, functions, or classes. When this happens, the interpreter creates the corresponding objects in memory and assigns them to the specified names — `def` and `class` are executable statements, not declarations. -5. The interpreter may also encounter statements that call functions or methods. When this happens, the interpreter looks up the function or method and executes the code it contains. -6. If the interpreter encounters an error while executing the code, it will raise an exception. If the error is not caught by the code, the interpreter will print a traceback and stop executing the program. +4. When the interpreter reaches statements that define variables, functions, or classes, it creates the corresponding objects in memory and binds them to their names — `def` and `class` are executable statements, not declarations. +5. When it reaches a function or method call, it looks up the function or method at that moment and executes its code. +6. If an error occurs while the code runs, the interpreter raises an exception. If nothing catches it, the interpreter prints a traceback and stops the program. -So Python is neither purely "interpreted" nor compiled to machine code: CPython compiles to bytecode and interprets that. (CPython 3.13+ additionally ships an experimental JIT, and alternative implementations like PyPy have JIT-compiled hot paths to machine code for years.) +So Python is neither purely "interpreted" nor compiled to machine code: CPython compiles to bytecode and interprets that. (Since 3.13, CPython can optionally be built with an experimental JIT compiler, and alternative implementations such as PyPy have compiled frequently executed code to machine code for years.) ## 86- What is `__pycache__`? -The `__pycache__` directory is a directory that is created by the Python interpreter to store compiled bytecode files. When you run a Python program, the interpreter converts the source code into a form that is more efficient to execute. This conversion process is known as compiling. The compiled bytecode files are stored in the `__pycache__` directory so that they can be used in future executions of the program without the need to recompile the source code. +`__pycache__` is a directory that the Python interpreter creates to store compiled bytecode (`.pyc`) files. Before running a module, Python compiles its source code into bytecode, a lower-level form that is faster to execute. Caching the bytecode on disk lets later runs skip this step, as long as the source file hasn't changed. Only imported modules are cached; the script you run directly is compiled fresh each time. -The `__pycache__` directory is typically located in the same directory as the Python source files that were used to create it. It is created automatically by the interpreter, and you do not need to worry about managing it manually. +The `__pycache__` directory is created next to the source files it caches. The interpreter manages it automatically: it is safe to delete, and it should be excluded from version control (for example, by adding it to `.gitignore`). You can stop Python from writing these files with the `-B` flag or the `PYTHONDONTWRITEBYTECODE` environment variable. -Note that the `__pycache__` directory and the compiled bytecode files it contains are specific to a particular version of Python. This means that if you change the version of Python that you are using, the interpreter will create a new `__pycache__` directory with compiled bytecode files that are compatible with the new version of Python. +Each cached file is tagged with the interpreter and version that produced it (for example, `module.cpython-312.pyc`). This means that different Python versions can share the same `__pycache__` directory without conflicts: each version simply writes and reads its own files. -## 87- What is the unittest in Python? +## 87- What is `unittest` in Python? -`unittest` is a unit testing framework in Python. It is a part of the Python Standard Library, and it is used to test small units of code, such as individual functions or methods. +`unittest` is the unit testing framework included in Python's standard library. It is used to test small units of code, such as individual functions or methods. -The `unittest` framework provides a set of tools for organizing and running tests, as well as for verifying the correctness of the code being tested. It includes a set of assertion methods that you can use to check the output of your code and ensure that it is correct. +The framework provides tools for organizing and running tests, plus a set of assertion methods for checking that your code produces the expected results. -To use the `unittest` framework, you define a series of test cases, which are individual units of testing that each test a specific aspect of your code. Each test case is a subclass of the `unittest.TestCase` class, and it includes a series of test methods that define the tests to be run. You can then use the `unittest` test runner to discover and run the tests in your test cases. +To use it, you write test cases as subclasses of `unittest.TestCase`. Each method whose name starts with `test` is a separate test. The `unittest` test runner then discovers and runs these tests and reports the results. Here is a simple example of how to use the `unittest` framework to test a function: @@ -3619,7 +3641,7 @@ if __name__ == '__main__': unittest.main() ``` -In this example, the `TestAdd` class defines two test methods, `test_add_two_positive_numbers` and `test_add_two_negative_numbers`, which test the `add` function with different input values. The `unittest.main()` function is used to run the tests and report the results. Test discovery is name-based: the runner executes methods whose names start with `test_`, and `python -m unittest discover` finds test modules across a project. +In this example, the `TestAdd` class defines two test methods, `test_add_two_positive_numbers` and `test_add_two_negative_numbers`, which test the `add` function with different input values. The `unittest.main()` function runs the tests and reports the results. Test discovery is name-based: the runner executes methods whose names start with `test`, and `python -m unittest discover` finds test modules (named `test*.py` by default) across a project. Beyond `assertEqual`, the pieces used daily are: `setUp`/`tearDown` (fresh fixtures before/after every test method), `assertRaises` as a context manager for expected exceptions, `assertAlmostEqual` for floats, and `unittest.mock` for patching out dependencies: @@ -3630,13 +3652,13 @@ class TestDivide(unittest.TestCase): 1 / 0 ``` -Worth saying in an interview: `unittest` is the standard library's xUnit-style framework, but much of the Python world uses **pytest**, which runs `unittest` suites unchanged while offering plain-`assert` tests, fixtures, and parametrisation with far less boilerplate. +Worth saying in an interview: `unittest` is the standard library's xUnit-style framework (modeled on Java's JUnit), but much of the Python world uses **pytest**, which runs `unittest` suites unchanged while offering plain-`assert` tests, fixtures, and parametrization with far less boilerplate (see question 140). -## 88- What is the difference between xrange and range? +## 88- What is the difference between `xrange` and `range`? -`range` and `xrange` are both functions that are used to generate a sequence of numbers. However, they differ in how they generate the numbers and in the type of object they return. +This is a Python 2 question: in Python 2, `range` and `xrange` both generate sequences of numbers, but they differ in how they generate them and in the type of object they return. -The `range` function generates a sequence of numbers by creating a list object that contains all of the numbers in the sequence. For example: +In Python 2, `range` builds a list that contains all of the numbers in the sequence. For example: ```python >>> range(5) @@ -3647,9 +3669,9 @@ The `range` function generates a sequence of numbers by creating a list object t [2, 4, 6, 8] ``` -The `range` function is useful when you need to generate a sequence of numbers and you need to access the numbers multiple times or perform operations on them. However, it can be inefficient when generating large sequences of numbers, as it creates a new list object in memory to hold the numbers. +A list is convenient when you need to reuse or modify the numbers, but it wastes memory for large ranges, because every number is stored up front. -The `xrange` function generates the numbers lazily instead, returning an `xrange` object rather than a `list`. It is a common mistake to call it a generator — it is a lazy **sequence**: it computes values on demand like a generator would, but it also supports `len()`, indexing, and repeated iteration, none of which a generator allows. The point stands that `xrange` is far more memory-efficient than Python 2's `range` for large spans, since it never materialises the whole sequence. +`xrange` generates the numbers lazily instead, returning an `xrange` object rather than a `list`. It is a common mistake to call it a generator — it is a lazy **sequence**: it computes values on demand like a generator would, but it also supports `len()`, indexing, and repeated iteration, none of which a generator allows. Either way, `xrange` is far more memory-efficient than Python 2's `range` for large ranges, because it never builds the whole list. For example: @@ -3668,7 +3690,7 @@ xrange(2, 10, 2) [2, 4, 6, 8] ``` -The `xrange` function was introduced in **`Python 2`** as a more efficient alternative to the `range`. In **`Python 3`**, the `range` function was redesigned along `xrange`'s lines, and `xrange` was removed from the language. +`xrange` was introduced in **Python 2** as a more efficient alternative to `range`. In **Python 3**, `range` was redesigned along the lines of `xrange`, and `xrange` was removed from the language. Python 3's `range` is in fact better than `xrange` ever was — it is a full lazy, immutable sequence: @@ -3682,7 +3704,7 @@ print(list(r[:5])) # [0, 1, 2, 3, 4] - slicing returns another range Because it is a sequence (not a generator), a `range` can be iterated any number of times, and its `in` test for integers is constant-time arithmetic rather than a scan. -## 89- What is the use of `//` operator in Python? +## 89- What is the `//` operator used for in Python? In Python, the `//` operator is the floor division operator: it divides and then applies `floor()`, rounding the result **down toward negative infinity**. @@ -3722,33 +3744,33 @@ Floor division pairs with the modulo operator through the invariant `a == (a // 1 ``` -## 90- How are dict and set implemented internally? What is the complexity of retrieving an item? How much memory do these structures consume? +## 90- How are `dict` and `set` implemented internally? What is the complexity of retrieving an item? How much memory do these structures consume? -In Python, both dictionaries (called `dict`) and `sets` are implemented using hash tables. A hash table is a data structure that uses a hash function to map keys to indices in an array, allowing for fast insertion, deletion, and lookup of keys. +In Python, both `dict` and `set` are implemented as hash tables. A hash table uses a hash function to turn each key into a position in an internal array, which allows fast insertion, deletion, and lookup. -In the case of `dict`, each key-value pair is stored in the hash table. The keys are used to calculate a hash value, which is used to determine the index in the array where the key-value pair should be stored. The value is then stored at that index. To retrieve a value from the `dict`, the hash function is used to calculate the index of the key-value pair, and the value is retrieved from that index. +For a `dict`, Python computes the hash of the key and uses it to pick the slot where the key and its value are stored. To look up a key, Python computes the hash again, goes straight to that slot, and confirms the match with `==`. A `set` works the same way, but stores only keys. -The complexity of retrieving an item from a `dict` or a `set` is typical `O(1)` on average, meaning that it takes a constant amount of time to retrieve an item, regardless of the size of the `dict` or `set`. However, in the worst case, the complexity can be `O(n)`, meaning that it takes linear time to retrieve an item if the hash function is poorly designed and causes many keys to hash to the same index. +Retrieving an item from a `dict` or a `set` takes **O(1)** time on average: the time does not grow with the size of the collection. In the worst case, it can degrade to **O(n)**, when many keys collide (produce the same slot), for example because of a poorly designed `__hash__` method. -As for memory consumption, the amount of memory that a `dict` or `set` consumes depends on the number of keys it contains and the size of the keys and values. In general, `dict` and `set` objects use more memory than other data structures, such as lists and tuples, because they store the keys and values in addition to the overhead of the hash table data structure. However, the exact amount of memory consumed will depend on the specific keys and values being stored and on the implementation of the Python interpreter. +As for memory, a `dict` or `set` uses noticeably more memory than a list or tuple with the same number of items. Besides storing the keys (and values), a hash table stores each key's hash and deliberately keeps part of its table empty so that lookups stay fast. The exact amount depends on the number of items and on the Python version. Implementation details worth knowing at a senior level: - **Collisions are handled by open addressing**, not chaining: on a collision, CPython probes other slots in the same table (with a perturbation scheme that mixes in more hash bits), rather than hanging linked lists off buckets. -- **The table resizes by load factor.** A dict keeps at most ~2/3 of its slots occupied; passing that threshold triggers a grow-and-rehash. This is why inserts are _amortised_ O(1) — an individual insert can pay for a full O(n) rehash. +- **The table resizes by load factor.** A dict keeps at most ~2/3 of its slots occupied; passing that threshold triggers a grow-and-rehash. This is why inserts are _amortized_ O(1): an individual insert may pay for a full O(n) rehash, but averaged over many inserts, the cost per insert stays constant. - **Modern dicts are "compact"** (CPython 3.6+): entries live in a dense array in insertion order, with the sparse hash table holding only small indices into it. This cut memory ~20-25% and is exactly why dicts preserve insertion order (guaranteed since 3.7). A `set` is essentially the same table without the values array — and sets do _not_ guarantee any order. - **Keys must be hashable** (immutable built-ins, or objects defining a consistent `__hash__`/`__eq__` pair). Mutable containers like lists are unhashable precisely because a mutated key could never be found again. - You can measure the overhead directly: `sys.getsizeof({})` versus `sys.getsizeof([])` shows the empty-container difference, and the gap grows with the sparse-slot overhead as items are added. -The practical consequence: membership tests are O(1) on dict/set versus O(n) on list/tuple, so the moment code does repeated `x in collection` checks over meaningful data sizes, converting the collection to a `set` is usually the single highest-value micro-optimisation available. +The practical consequence: membership tests are O(1) on dict/set versus O(n) on list/tuple, so the moment code does repeated `x in collection` checks over meaningful data sizes, converting the collection to a `set` is usually the single most valuable micro-optimization available. -## 91- What is MRO in Python? How does it work? +## 91- What is the MRO in Python, and how does it work? -In Python, MRO stands for "Method Resolution Order." It is a mechanism that is used to determine the order in which the methods of a class should be inherited when a class is derived from multiple base classes. +In Python, MRO stands for "Method Resolution Order." It is the order in which Python searches a class and its base classes when looking up a method or attribute. -In Python, a class can be derived from multiple base classes, creating a class hierarchy. When a class is derived from multiple base classes, it is said to have multiple inheritance. In multiple inheritance, a class can inherit methods from multiple base classes, and it is important to determine the order in which these methods should be inherited to avoid conflicts. +The MRO matters most with multiple inheritance, where a class inherits from several base classes and the same method name may exist in more than one of them. The MRO decides which version wins. -The MRO of a class is the order in which the methods of the class and its base classes are searched when looking up a method. In Python, the MRO of a class is determined using the C3 linearization algorithm, which produces a linear order that preserves the local precedence order of the base classes. +Python computes the MRO with the C3 linearization algorithm, which produces a single ordered list of classes that respects the order in which base classes are listed in each class definition. Here is an example of a class hierarchy with multiple inheritance in Python: @@ -3775,13 +3797,11 @@ In this example, the class `D` is derived from the classes `B`, `C`, and `A`, in D.__mro__ == (D, B, C, A, object) ``` -This means that when looking up a method on an instance of the `D` class, the interpreter will first search the `D` class, then the `B` class, then the `C` class, then the `A` class, and finally the object class, which is the base class of all classes in Python. - -The MRO is an important concept in Python because it determines the order in which methods are inherited and how conflicts are resolved when a class has multiple inheritance. Understanding how the MRO works is essential to understanding how multiple inheritance works in Python. +This means that when looking up a method on an instance of `D`, the interpreter searches `D` first, then `B`, then `C`, then `A`, and finally `object`, the base class of all classes in Python. So `D().foo()` prints `B.foo`. -Three follow-ups an interviewer is likely to reach for: +Three follow-ups an interviewer is likely to ask about: -- **C3 is more than left-to-right.** The linearization satisfies two constraints at once: a class always precedes its own bases, and the left-to-right order of the bases listed in every class definition is preserved. In diamond hierarchies this means a shared base appears _after_ all its subclasses, not immediately after the first parent (see the `super` question earlier in this file for a worked diamond). +- **C3 is more than left-to-right.** The linearization satisfies two constraints at once: a class always precedes its own bases, and the left-to-right order of the bases listed in every class definition is preserved. In diamond hierarchies, this means a shared base appears _after_ all its subclasses, not immediately after the first parent (see question 31 for a worked diamond example). - **Not every hierarchy has a valid MRO.** If the constraints contradict each other, Python refuses to create the class at all: ```python @@ -3792,9 +3812,9 @@ Three follow-ups an interviewer is likely to reach for: - **`super()` is MRO traversal.** `super()` does not mean "my parent" — it means "the next class after mine in the MRO of the instance's actual type". That is what allows cooperative multiple inheritance and mixins to compose: each class calls `super()`, and the MRO threads one call through every class exactly once. Inspect it any time with `D.__mro__` or `D.mro()`. -## 92- How to distribute Python code? +## 92- How do you distribute Python code? -There are several ways to distribute Python code, depending on the specific needs of your project. Here are a few common options: +There are several ways to distribute Python code, depending on the needs of your project. Here are the most common options: 1. **Packaging for PyPI (the modern workflow)**: Declare the package metadata in a `pyproject.toml` file — this has replaced the old `setup.py`/`distutils` approach (`distutils` was removed from the standard library in Python 3.12). Then build and upload: @@ -3803,51 +3823,51 @@ There are several ways to distribute Python code, depending on the specific need python -m twine upload dist/* ``` - Users then install it with `pip install your-package`. Tools like Poetry, Hatch, and uv wrap this same standards-based flow (PEP 517/518/621) with dependency management on top. `setuptools` still works fine as the build backend — but configured via `pyproject.toml`, with `setup.py` kept only for legacy or complex native builds. + Users then install it with `pip install your-package`. Tools like Poetry, Hatch, and uv wrap this same standards-based flow (PEP 517/518/621) and add dependency management on top. `setuptools` still works fine as the build backend — but configured via `pyproject.toml`, with `setup.py` kept only for legacy projects or complex native builds. Question 150 walks through the publishing steps in detail. -2. **Distributing within a team without PyPI**: `pip` can install straight from a git URL (`pip install git+https://github.com/org/repo.git`), from a private index (`--index-url`, e.g. an internal devpi/Artifactory), or from a local wheel file. Wheels are the unit of distribution either way (see the wheels-vs-eggs question earlier). +2. **Distributing within a team without PyPI**: `pip` can install straight from a Git URL (`pip install git+https://github.com/org/repo.git`), from a private package index (`--index-url`, e.g., an internal devpi or Artifactory server), or from a local wheel file. Wheels are the unit of distribution either way (see question 41). -3. **Distributing as a standalone executable**: To ship to users who do not have Python installed, bundle the interpreter and dependencies into one artifact with **PyInstaller** (the de facto standard, cross-platform), or alternatives like cx_Freeze, Briefcase (GUI apps), Nuitka (compiles to C), or `shiv`/`pex` (self-contained zipapps building on the stdlib `zipapp` module). The older `py2exe`/`py2app` tools fill the same niche but are platform-specific. +3. **Distributing as a standalone executable**: To ship to users who do not have Python installed, bundle the interpreter and all dependencies into a single artifact with **PyInstaller** (the de facto standard, cross-platform), or alternatives such as cx_Freeze, Briefcase (GUI apps), Nuitka (compiles to C), or `shiv`/`pex` (self-contained zip applications, built on the standard library's `zipapp` module). The older `py2exe` and `py2app` tools serve the same purpose but are tied to one platform each. -4. **Containers**: For services, the answer in practice is often a Docker image — the dependency story (OS libraries included) is pinned once and runs anywhere a container runtime exists. +4. **Containers**: For services, the practical answer is often a Docker image — all dependencies, including OS libraries, are pinned once, and the image runs anywhere a container runtime is available. -## 93- How to work with Python transitive dependencies? +## 93- How do you manage transitive dependencies in Python? -Transitive dependencies are dependencies that are required by a package that your code depends on. For example, if your code depends on the package `A`, and package `A` depends on package `B`, then package `B` is a transitive dependency of your code. +Transitive dependencies are the dependencies of your dependencies. For example, if your code depends on package `A`, and package `A` depends on package `B`, then package `B` is a transitive dependency of your code. -To work with transitive dependencies in Python, you typically use a package manager like `pip` to install and manage your dependencies. When you install a package using `pip`, it will automatically install any transitive dependencies that the package requires. +In Python, you typically manage dependencies with a package installer such as `pip`. When you install a package with `pip`, it automatically installs every transitive dependency that the package requires. -For example, suppose you have a Python project that depends on package `A`, which in turn depends on package `B`. To install these dependencies using `pip`, you can use the following command: +For example, suppose your project depends on package `A`, which in turn depends on package `B`. To install them, run: ```bash pip install A ``` -This will install both packages `A` and its transitive dependency, package `B`. +This installs both package `A` and its transitive dependency, package `B`. -If you want to specify the exact version of a package and its transitive dependencies that you want to install, you can use the `-r` flag to specify a requirements file. A `requirements` file is a text file that lists the packages and their versions that your project depends on. For example: +To install exact versions of a package and its transitive dependencies, list them in a requirements file and pass it to `pip` with the `-r` flag. A requirements file is a text file that lists the packages your project depends on, with their versions. For example: -```python +```text A==1.0 B==2.0 ``` -To install the packages and their transitive dependencies from this requirements file, you can use the following command: +To install the packages from this requirements file, run: ```bash pip install -r requirements.txt ``` -This will install package `A` version **`1.0`** and its transitive dependency, package `B` version **`2.0`**. +This installs package `A` version **`1.0`** and its transitive dependency, package `B` version **`2.0`**. -Using a package manager like `pip` to manage your transitive dependencies is a good way to ensure that your code has the correct dependencies installed and to keep them up to date. It also makes it easier to share your code with others, as they can use the requirements file to install the correct dependencies for your project. +Pinning versions this way ensures that everyone who installs your project — teammates, CI, and production — gets the same set of dependencies. The senior-level practice is to separate **direct** dependencies from the **fully pinned** set: - Declare only your direct dependencies with loose constraints (in `pyproject.toml`, or a hand-edited `requirements.in`). - Generate a **lock file** pinning every transitive dependency to an exact version for reproducible deploys — `pip freeze > requirements.txt` is the crude form; `pip-compile` (pip-tools), Poetry, or uv produce proper lock files with hashes. -- `pip install` resolves the whole graph and refuses genuinely conflicting constraints (since the 2020 resolver); `pip check` verifies an existing environment, and `pipdeptree` visualises who pulls in what — the first tool to reach for when a mystery package appears in your environment. -- Do all of this inside a **virtual environment** (`python -m venv`), so project graphs cannot contaminate each other or the system Python. +- `pip install` resolves the whole dependency tree and refuses truly conflicting version requirements (since its new resolver in 2020); `pip check` verifies an existing environment, and `pipdeptree` shows which package pulls in which — the first tool to reach for when an unexpected package appears in your environment. +- Do all of this inside a **virtual environment** (`python -m venv`), so the dependencies of different projects cannot conflict with each other or with the system Python. ## 94- What is the output of this code? @@ -3874,7 +3894,7 @@ except StopIteration as e: print(e.value) # 666 ``` -A `for` loop swallows the `StopIteration` silently, so the `666` is invisible to ordinary iteration. The mechanism exists for **generator delegation**: inside another generator, `result = yield from Foo()` re-yields the `42` and binds `result` to `666`. This is the foundation coroutines were originally built on, and it is why "what does `return` do in a generator?" is a favourite senior-level probe. (The trailing semicolon after `yield 42;` is legal but un-Pythonic.) +A `for` loop swallows the `StopIteration` silently, so the `666` is invisible to ordinary iteration. The mechanism exists for **generator delegation**: inside another generator, `result = yield from Foo()` re-yields the `42` and binds `result` to `666`. This is the foundation coroutines were originally built on (see question 154), and it is why "what does `return` do in a generator?" is a favorite senior-level interview question. (The trailing semicolon after `yield 42;` is legal but un-Pythonic.) ## 95- What is the output of this code? @@ -3886,9 +3906,9 @@ class MangledGlobal: return __mangled ``` -The `MangledGlobal` class contains a reference to a global variable with a name that has been "`mangled`" to avoid name conflicts with other variables in the global namespace. +Strictly speaking, this code produces no output — it only defines a class. The real question is what `MangledGlobal().test()` returns, and the answer is `23`. -In Python, the name `mangling` is a technique that is used to protect instance variables in a class from being accidentally overwritten by derived classes. Name `mangling` works by adding a double underscore prefix to the name of an instance variable, which causes the interpreter to automatically rename the variable in a way that is unique to the class. +The `test` method refers to `__mangled`, a name that Python _mangles_ (rewrites) because it starts with two underscores. Name mangling is a technique that protects a class's attributes from being accidentally overridden by subclasses: inside a class body, an identifier such as `__name` is renamed to `_ClassName__name`, which makes it unique to that class. The detail this puzzle turns on: mangling is applied **at compile time to any identifier of the form `__name` appearing anywhere in a class body** — not just to attribute access through `self`. So the bare reference `__mangled` inside `test()` is textually rewritten to `_MangledGlobal__mangled` before the code ever runs. At call time, ordinary name lookup then proceeds (local → global): there is no local by that name, but the _global_ `_MangledGlobal__mangled = 23` matches the rewritten name, so `test()` returns `23`. @@ -3910,40 +3930,38 @@ print(_MangledGlobal__mangled) # Output: 23 Mangling applies only to identifiers with **two leading underscores and at most one trailing underscore** — so `__x` is mangled, while `__x__` (dunders) and `_x` are not. Its intended purpose is to keep a base class's private attributes from being accidentally overridden in subclasses, since each class mangles to its own name. -Note that while the name `mangling` is intended to protect instance variables from being overwritten by derived classes, it is not a security feature and should not be relied upon to protect sensitive data. The name `mangling` can be easily bypassed by using the mangled name directly, as shown in the example above. +Note that name mangling is not a security feature and should not be relied on to protect sensitive data: anyone can still access the attribute through its mangled name, as the example above shows. -## 96- What is packing and unpacking in Python? +## 96- What are packing and unpacking in Python? -In Python, packing and unpacking refer to two related concepts involving the conversion of data between different structures. +Packing and unpacking are two sides of the same idea: packing combines several values into one collection, and unpacking splits a collection back into separate values. -Packing refers to the process of taking multiple values or items and combining them into a single data structure. This is often done using tuples, which are a type of immutable data structure in Python. For example, we can create a tuple that contains three elements like this: +The most common example of packing is a tuple: writing values separated by commas packs them into a tuple (the parentheses are optional). For example: ```python my_tuple = (1, "hello", True) ``` -Unpacking, on the other hand, refers to the process of taking a data structure and splitting it into multiple values or items. This is often done using tuples, lists, or dictionaries. For example, we can unpack a tuple into multiple variables like this: +Unpacking does the reverse: it assigns the elements of a collection to separate variables. Any iterable can be unpacked, as long as the number of variables matches the number of values. For example, we can unpack a tuple into three variables like this: ```python my_tuple = (1, "hello", True) a, b, c = my_tuple ``` -Also, the asterisk (\*) symbol can be used in unpacking expressions to represent a variable number of elements. This is sometimes referred to as "extended unpacking". - -The asterisk can be used in several ways: +The asterisk (`*`) can also be used in unpacking to stand for a variable number of elements. This is known as "extended unpacking," and it can be used in several ways: -1. Unpacking into individual variables: If you have a list or tuple with an unknown number of elements, you can use the asterisk to unpack the elements into individual variables. For example: +1. Collecting the remaining elements: If a list or tuple has an unknown number of elements, a starred variable collects all the leftover elements into a list. For example: ```python my_list = [1, 2, 3, 4, 5] a, b, *rest = my_list - print(a) # 1 - print(b) # 2 - print(rest) # [3, 4, 5] + print(a) # 1 + print(b) # 2 + print(rest) # [3, 4, 5] ``` -2. Unpacking in function calls: The asterisk can also be used to unpack arguments in function calls. For example: +2. Unpacking in function calls: The asterisk can also be used to spread a collection into separate arguments in a function call. For example: ```python def my_function(a, b, c): @@ -3955,7 +3973,7 @@ The asterisk can be used in several ways: In this example, the elements of `my_list` are unpacked and passed as arguments to the function `my_function`. -The picture is completed by the double asterisk and the packing side of function signatures: +The double asterisk and the packing side of function signatures complete the picture: 1. **Packing in function signatures**: `*args` packs surplus positional arguments into a tuple and `**kwargs` packs surplus keyword arguments into a dict — packing and unpacking are the same syntax viewed from opposite ends: @@ -3979,17 +3997,17 @@ The picture is completed by the double asterisk and the packing side of function combined = [*range(3), *"ab"] # [0, 1, 2, 'a', 'b'] ``` -3. **Swap and starred assignment**: the idiomatic `a, b = b, a` is packing and unpacking in one statement — the right side packs into a tuple, the left side unpacks it. Unpacking also works in `for` loops over pairs (`for key, value in d.items():`). +3. **Swapping and loops**: the idiomatic `a, b = b, a` is packing and unpacking in one statement — the right side packs into a tuple, the left side unpacks it. Unpacking also works in `for` loops over pairs (`for key, value in d.items():`). ## 97- What's the difference between `globals()`, `locals()`, and `vars()`? -In Python, the `globals()`, `locals()`, and `vars()` functions are `built-in` functions that can be used to retrieve the global, local, and instance variables in a program, respectively. +In Python, `globals()`, `locals()`, and `vars()` are built-in functions that return namespaces as dictionaries: roughly, the global variables, the local variables, and an object's attributes, respectively. -The `globals()` function returns a dictionary that contains the `global` variables in the current program. **Global variables** are variables that are defined at the top level of a module or script and are accessible from anywhere in the program. +The `globals()` function returns a dictionary that contains the global variables of the current module. **Global variables** are variables that are defined at the top level of a module or script and are accessible from anywhere in that module. The `locals()` function returns a dictionary that contains the local variables in the current function or method. **Local variables** are variables that are defined within a function or method and are only accessible within that function or method. -The `vars()` function has two distinct behaviours. **Without arguments, `vars()` is equivalent to `locals()`** — it returns the current local namespace. **With an argument, `vars(obj)` returns `obj.__dict__`** — the attribute dictionary of a module, class, or instance. It is the "with an argument" form that makes `vars()` interesting: `vars(some_instance)` shows an object's instance attributes, which neither `globals()` nor `locals()` can do. (An object with no `__dict__` — for example one using `__slots__`, or a plain `int` — makes `vars(obj)` raise `TypeError`.) +The `vars()` function has two distinct behaviors. **Without arguments, `vars()` is equivalent to `locals()`** — it returns the current local namespace. **With an argument, `vars(obj)` returns `obj.__dict__`** — the attribute dictionary of a module, class, or instance. It is the "with an argument" form that makes `vars()` interesting: `vars(some_instance)` shows an object's instance attributes, which neither `globals()` nor `locals()` can do. (An object with no `__dict__` — for example, one using `__slots__`, or a plain `int` — makes `vars(obj)` raise `TypeError`.) Here is an example of how you might use these functions: @@ -4025,14 +4043,14 @@ vars(Point(1, 2)): {'px': 1, 'py': 2} Two cautions that matter in practice: -- **Writing to `globals()` works but is a design smell; writing to `locals()` inside a function does not work at all.** Function locals are stored in optimised slots, and (in 3.13+, per PEP 667) `locals()` returns a _snapshot_ — mutating the returned dict never changes the actual variables. -- These tools are for debugging and framework plumbing (e.g. `str.format_map(vars(obj))`); reaching for them in ordinary application logic usually signals that a plain dict should have been used instead. +- **Writing to `globals()` works but is usually a sign of poor design; writing to `locals()` inside a function does not work at all.** A function's local variables are stored in fixed slots rather than in a dictionary, and (in 3.13+, per PEP 667) `locals()` returns a _snapshot_ — changing the returned dict never changes the actual variables. +- These tools are meant for debugging and framework internals (e.g., `template.format_map(vars(obj))`); needing them in ordinary application logic usually means a plain dict should have been used instead. -## 98- What is the `__init__.py` module, and what is it for? +## 98- What is the `__init__.py` file, and what is it for? -The `__init__.py` file is used to mark directories on disk as Python package directories. It is required for Python to treat the directories as containing packages; otherwise, the directories are just treated as directories and are not searched for modules. +The `__init__.py` file marks a directory as a regular Python package. Traditionally, it was required: without it, Python would not import modules from the directory (this changed in Python 3.3; see the nuance below). -The `__init__.py` file can contain code that initializes the package or sets up any additional functionality that the package provides. It is executed when the package is imported. +The `__init__.py` file can also contain code that initializes the package or sets up functionality the package provides. This code runs when the package is first imported. For example, consider the following directory structure: @@ -4052,15 +4070,15 @@ To import `module1` from the `my_package` package, you would use the following i import my_package.module1 ``` -When this `import` statement is executed, Python will execute the code in `my_package/__init__.py` before it loads `module1`. +When this `import` statement is executed, Python runs the code in `my_package/__init__.py` before it loads `module1`. -The `__init__.py` file can be an empty file, and typical non-empty uses are: re-exporting the package's public API so users can write `from my_package import Thing` instead of digging into submodules, defining `__all__`, and setting package-level metadata. +The `__init__.py` file is often empty. When it isn't, its typical uses are: re-exporting the package's public API so users can write `from my_package import Thing` instead of digging into submodules, defining `__all__`, and setting package-level metadata. -The modern nuance: since Python 3.3 (PEP 420), a directory **without** `__init__.py` still imports — it becomes an implicit _namespace package_, whose parts can even be spread across multiple `sys.path` locations. So "it must be present" is no longer strictly true. In practice you should still add `__init__.py` to every ordinary package: it makes the package explicit, imports marginally faster, plays better with some tools, and prevents two unrelated directories from silently merging into one package — reserving namespace packages for the rare plugin-style layouts that genuinely need them. +The modern nuance: since Python 3.3 (PEP 420), a directory **without** `__init__.py` still imports — it becomes an implicit _namespace package_, whose parts can even be spread across multiple `sys.path` locations. So "it must be present" is no longer strictly true. In practice, you should still add `__init__.py` to every ordinary package: it makes the package explicit, imports slightly faster, works better with some tools, and prevents two unrelated directories from silently merging into one package. Reserve namespace packages for the rare plugin-style layouts that genuinely need them. -## 99- How do I view object methods? +## 99- How do you view an object's methods? -To view the methods of an object in Python, you can use the `dir()` function. This function returns a `list` of all the attributes and methods of an object, including special attributes like `__dict__` and `__doc__`. +To view the methods of an object in Python, you can use the `dir()` function. It returns a sorted list of the names of an object's attributes and methods, including special ones like `__dict__` and `__doc__`. For example, consider the following object: @@ -4081,10 +4099,10 @@ methods = dir(obj) print(methods) ``` -This would output the following list: +This would output the following list (on Python 3.12; the exact list varies slightly between versions): -```bash -['__class__', '__delattr__', '__dict__', '__dir__', '__doc__', '__eq__', '__format__', '__ge__', '__getattribute__', '__gt__', '__hash__', '__init__', +```text +['__class__', '__delattr__', '__dict__', '__dir__', '__doc__', '__eq__', '__format__', '__ge__', '__getattribute__', '__getstate__', '__gt__', '__hash__', '__init__', '__init_subclass__', '__le__', '__lt__', '__module__', '__ne__', '__new__', '__reduce__', '__reduce_ex__', '__repr__', '__setattr__', '__sizeof__', '__str__', '__subclasshook__', '__weakref__', 'my_method', 'x'] ``` @@ -4098,30 +4116,30 @@ print(obj_methods) This would output the following list: -```python +```text ['my_method'] ``` -Other tools for the same job: `help(obj)` renders methods with their signatures and docstrings; `inspect.getmembers(obj, inspect.ismethod)` returns `(name, method)` pairs and is the robust programmatic option (`inspect.isfunction` for the unbound functions on the class itself); and `vars(type(obj))` shows what the class defines directly, excluding what it merely inherits. One caution: `dir()` calls are a _convention_, not a guarantee — a class can override `__dir__`, and dynamic attributes served by `__getattr__` will not appear. +Other tools for the same job: `help(obj)` renders methods with their signatures and docstrings; `inspect.getmembers(obj, inspect.ismethod)` returns `(name, method)` pairs and is the robust programmatic option (use `inspect.isfunction` for the plain functions defined on the class itself); and `vars(type(obj))` shows what the class defines directly, excluding what it inherits. One caution: `dir()` is a best-effort listing, not a guarantee — a class can override `__dir__`, and dynamic attributes provided by `__getattr__` will not appear. -## 100- Which is a better practice - global import or local import in Python +## 100- Which is better practice in Python: global imports or local imports? -The convention — codified in PEP 8 — is the opposite of what this question often tempts people to say: **imports belong at the top of the module** (global imports). Top-level imports make a module's dependencies visible at a glance, fail fast at import time rather than deep inside a call at 3 a.m., and cost nothing on reuse — Python caches every imported module in `sys.modules`, so repeated imports are just a dictionary hit, and a function-local import actually _adds_ a small lookup cost on every call. +The convention, set out in PEP 8, is clear: **imports belong at the top of the module** (global imports). Top-level imports make a module's dependencies visible at a glance, fail immediately at import time rather than deep inside a call at 3 a.m., and cost nothing on reuse. Python caches every imported module in `sys.modules`, so repeated imports are just a dictionary lookup — and a function-local import actually _adds_ a small lookup cost on every call. Local (function-level) imports are the exception, justified in a few specific situations: -1. **Breaking a circular import**, when two modules genuinely need each other and restructuring is not yet feasible — moving one import inside a function defers it past module initialisation. -2. **Deferring a heavy dependency** to speed up program start-up: if `import pandas` costs seconds and only one rarely-used command needs it, importing it inside that function keeps the CLI snappy. -3. **Optional dependencies**: a feature that needs an extras-installed package imports it locally (typically inside a `try`/`except ImportError`) so the rest of the module works without it. +1. **Breaking a circular import**, when two modules genuinely need each other and restructuring is not yet feasible — moving one import inside a function delays it until after both modules have finished initializing. +2. **Deferring a heavy dependency** to speed up program startup: if `import pandas` takes seconds and only one rarely used command needs it, importing it inside that function keeps the command-line tool fast to start. +3. **Optional dependencies**: a feature that needs a package installed only as an optional extra imports it locally (typically inside a `try`/`except ImportError`), so the rest of the module works without it. 4. **Platform- or context-specific modules** that may not exist everywhere (`fcntl` on Windows, test-only helpers). Note that "avoiding name conflicts" is _not_ a good reason — aliasing handles that at the top of the file (`import json as std_json`). The practical rule: top-level imports by default; a local import is a deliberate, commented exception, not a style choice. -## 101- what is tilde symbol `(~)` used for in Python? +## 101- What is the tilde symbol (`~`) used for in Python? -In a `requirements` file for Python packages, the tilde symbol `(~)` is used to specify a version constraint. For example, if a package `foo` requires version `bar` equal to or greater than `1.2.3` but less than `1.3.0`, the constraint could be specified as `foo~=1.2.3`. +In a requirements file (and anywhere else `pip` accepts version specifiers), `~=` is the _compatible release_ operator defined by PEP 440. For example, `foo~=1.2.3` means "version `1.2.3` or later, but below `1.3.0`." -This syntax is used to specify a minimum version of the package, as well as allow for updates that might be made to the package that is compatible with the project's requirements. It allows for patch-level updates (e.g. `1.2.3` to `1.2.4`) but not for updates that might introduce backward-incompatible changes (e.g. `1.2.3` to `1.3.0`). +It sets a minimum version while still allowing compatible updates: here, patch-level releases (e.g., `1.2.3` to `1.2.4`), but not releases that might introduce breaking changes (e.g., `1.3.0`). With only two version components, `foo~=1.2` means `>=1.2, <2.0`. For example, a `requirements` file might contain the following line: @@ -4131,7 +4149,7 @@ foo~=1.2.3 This would install the latest version of `foo` that is equal to or greater than `1.2.3` and less than `1.3.0`. -Also, the tilde symbol `(~)` is the bitwise `NOT` operator (implemented by the `__invert__` dunder). On an integer it inverts all bits, which under two's-complement arithmetic means **`~x` equals `-x - 1`**: +In Python code itself, `~` is the bitwise `NOT` operator (implemented by the `__invert__` dunder method). On an integer, it inverts all bits, which under two's-complement arithmetic (the way integers are represented in binary) means **`~x` equals `-x - 1`**: ```python >>> x = 0b1100 # 12 @@ -4146,7 +4164,7 @@ Also, the tilde symbol `(~)` is the bitwise `NOT` operator (implemented by the ` In this example, the value of `x` is `12` (`1100` in binary), and `~x` is `-13` (the REPL displays the integer `-13`; `bin()` shows its binary form). -The `-x - 1` identity gives rise to a neat indexing idiom: `~i` mirrors an index from the other end of a sequence, since `lst[~0]` is the last element, `lst[~1]` the second-to-last, and in general `lst[~i] == lst[-i - 1]`. It also makes `~` the go-to _element-wise_ NOT in the scientific stack — in NumPy and pandas, `df[~mask]` selects the rows where a boolean mask is `False` (plain `not` cannot be overloaded for arrays). On Python's own `bool` values, though, beware: `~True` is `-2`, not `False` — use `not` for scalar logic. +The `-x - 1` identity gives rise to a neat indexing idiom: `~i` mirrors an index from the other end of a sequence, since `lst[~0]` is the last element, `lst[~1]` the second-to-last, and in general `lst[~i] == lst[-i - 1]`. In NumPy and pandas, `~` is also the standard _element-wise_ NOT: `df[~mask]` selects the rows where a boolean mask is `False` (plain `not` cannot be overloaded for arrays). On Python's own `bool` values, though, beware: `~True` is `-2`, not `False` — use `not` for scalar logic. ## 102- What is the difference between `__str__` and `__repr__`? @@ -4154,7 +4172,7 @@ Both `__str__` and `__repr__` are special methods in Python that can be used to 1. `__str__` is used to return a human-readable string representation of an object. It is intended to be used for display purposes, such as printing the object to the console or displaying it in a GUI. `__str__` should be easy to read and understand for humans, and it should not contain unnecessary technical details. -2. `__repr__` is used to return an unambiguous string representation of an object. It is intended to be used for debugging purposes, such as inspecting the object's state or reproducing the object in code. `__repr__` should be a valid Python expression that can be used to recreate the object, and it should contain all the relevant technical details about the object. +2. `__repr__` is used to return an unambiguous string representation of an object. It is intended for developers, for debugging and logging, such as inspecting the object's state. Ideally, `__repr__` returns a valid Python expression that would recreate the object; when that isn't practical, the convention is a description in angle brackets, such as ``. In summary, `__str__` is used for display purposes and should be easy to read, while `__repr__` is used for debugging purposes and should contain all the relevant technical details. @@ -4185,15 +4203,15 @@ print(str(today), "|", repr(today)) # 2026-07-25 | datetime.date(2026, 7, 25) <- the stdlib models the distinction ``` -Note the container behaviour: printing a list of objects shows each element's `__repr__`, never its `__str__` — a frequent source of "why is my nice string not showing?" confusion. In f-strings, `!r` requests the repr explicitly. `dataclasses.dataclass` generates a sensible `__repr__` for free, which is one more reason to reach for it for data-holding classes. +Note the container behavior: printing a list of objects shows each element's `__repr__`, never its `__str__` — a frequent source of "why is my nice string not showing?" confusion. In f-strings, `!r` requests the repr explicitly. `dataclasses.dataclass` generates a sensible `__repr__` for free, which is one more reason to use it for classes that mainly hold data. -## 103- What is `lru_cache` decorator in Python? +## 103- What is the `lru_cache` decorator in Python? -`lru_cache` is a decorator provided by the Python Standard Library's functools module that is used to cache the results of a function. It stands for "Least Recently Used Cache" and is a technique that can be used to speed up the execution of a function by caching its results. +`lru_cache` is a decorator from the standard library's `functools` module that caches the results of a function. LRU stands for "least recently used": when the cache is full, the entry that was used least recently is discarded to make room for a new one. -The `lru_cache` decorator works by storing the results of the function in a cache dictionary. If the same set of arguments is passed to the function again, the cached result is returned instead of executing the function again. This can be useful when the function takes a long time to execute, or when the same set of arguments are used repeatedly. +The decorator stores the function's results in a dictionary, keyed by the arguments. If the same arguments are passed again, the cached result is returned instead of running the function again. This is useful when the function is slow and is called repeatedly with the same arguments. -The `lru_cache` decorator has several optional arguments, such as `maxsize` which determines the maximum number of results to cache, and `typed` which determines whether to treat arguments of different types separately. By default, the `lru_cache` decorator uses a maximum cache size of `128`. +The decorator has two optional arguments: `maxsize`, which sets the maximum number of results to cache (`128` by default), and `typed`, which determines whether arguments of different types are cached separately. Here's an example of how to use the `lru_cache` decorator in Python: @@ -4205,12 +4223,12 @@ def fibonacci(n): if n < 2: return n else: - return fibonacci(n-1) + fibonacci(n-2) + return fibonacci(n - 1) + fibonacci(n - 2) print(fibonacci(30)) ``` -In this example, the memoisation transforms the algorithm's complexity: naive recursive Fibonacci recomputes the same subproblems exponentially many times — O(2n) — while with `lru_cache` each `fibonacci(k)` is computed once and every repeat becomes a cache hit, collapsing the cost to O(n). A second call to `fibonacci(30)` doesn't even recurse; it is a single dictionary lookup. +In this example, memoization (reusing the saved results of earlier calls) transforms the algorithm's complexity: naive recursive Fibonacci recomputes the same subproblems exponentially many times — O(2n) — while with `lru_cache` each `fibonacci(k)` is computed once and every repeat becomes a cache hit, collapsing the cost to O(n). A second call to `fibonacci(30)` doesn't even recurse; it is a single dictionary lookup. Operational details a senior engineer should have ready: @@ -4218,15 +4236,15 @@ Operational details a senior engineer should have ready: - **Inspect and reset** with `fibonacci.cache_info()` (hits, misses, maxsize, currsize) and `fibonacci.cache_clear()` — `cache_info()` is also handy in tests to assert caching actually happens. - `maxsize=None` means an unbounded cache with no LRU eviction bookkeeping; `functools.cache` (Python 3.9+) is a clearer alias for exactly that. `typed=True` caches `f(1)` and `f(1.0)` separately. - **Two classic traps**: decorating a _method_ keeps `self` in every cache key, so instances are never garbage-collected while cached (prefer `functools.cached_property`, or a cache scoped inside the instance); and caching a function that returns a **mutable** object hands every caller the same instance — mutate it and you have poisoned the cache. -- The cache is per-process and thread-safe for CPython's purposes, but it is not shared across processes — web workers each warm their own. +- The cache is thread-safe (concurrent calls won't corrupt it, although two threads may occasionally compute the same value at the same time), but it lives inside one process and is not shared — each web server worker process builds its own. ## 104- What does `__all__` do? -In Python, `__all__` is a special variable that can be defined at the top of a module, which is a list of strings that defines what symbols (e.g., functions, classes, and variables) the module exports when other modules import it using the `from module import *` syntax. +In Python, `__all__` is a module-level list of strings that names the public symbols (e.g., functions, classes, and variables) a module exports when another module uses `from module import *`. -When a module is imported using the `from module import *` syntax, Python only imports the names listed in the module's `__all__` list (if it is defined). If `__all__` is not defined, Python will import all names that do not start with an underscore (_). +When a module is imported with `from module import *`, Python imports only the names listed in the module's `__all__` (if it is defined). If `__all__` is not defined, Python imports all names that do not start with an underscore (`_`). -Defining `__all__` can be useful for controlling the public interface of a module, especially for large modules with many symbols. By specifying a list of only the public symbols, the module author can prevent accidental imports of internal or private symbols, which can reduce naming conflicts and make the code easier to understand. +Defining `__all__` is useful for controlling the public interface of a module, especially a large module with many symbols. By listing only the public symbols, the module author prevents internal or private names from leaking into other modules, which reduces naming conflicts and makes the code easier to understand. Here's an example of how `__all__` can be used: @@ -4242,55 +4260,57 @@ def _bar(): __all__ = ['foo'] ``` -In this example, `my_module` defines two functions foo and `_bar`. `_bar` is intended to be used only within the module and is not meant to be part of the module's public API. By setting `__all__` to `['foo']`, we tell Python to only export the foo function when other modules import `my_module` using the `from my_module import *` syntax. +In this example, `my_module` defines two functions, `foo` and `_bar`. `_bar` is intended to be used only within the module and is not part of the module's public API. By setting `__all__` to `['foo']`, we tell Python to export only `foo` when another module uses `from my_module import *`. + +Note that `__all__` affects only `import *`: an explicit import such as `from my_module import _bar` still works. It also documents the public API for readers and for tools such as linters, IDEs, and documentation generators. ## 105- List some of the dunder methods -1. `__init__`: Initializer, called on a freshly created instance to set it up (the object is actually _created_ by `__new__`, which is the rarely-overridden true constructor) -2. `__repr__`: Method that returns a printable representation of an object -3. `__str__`: Method that returns a string representation of an object -4. `__len__`: Method that returns the length of an object -5. `__getitem__`: Method that allows you to access an item in an object using the square bracket notation (`[]`) -6. `__setitem__`: Method that allows you to set an item in an object using the square bracket notation (`[]`) -7. `__delitem__`: Method that allows you to delete an item from an object using the square bracket notation (`[]`) -8. `__iter__`: Method that returns an iterator for an object -9. `__next__`: Method that returns the next value from an iterator -10. `__call__`: Method that allows you to call an object as if it were a function -11. `__getattr__`: Method that is called when an attribute is not found in an object -12. `__setattr__`: Method that is called when an attribute is set in an object -13. `__delattr__`: Method that is called when an attribute is deleted from an object -14. `__enter__`: Method that is called when a context manager is entered -15. `__exit__`: Method that is called when a context manager is exited -16. `__hash__`: Method that returns a hash value for an object -17. `__bool__`: Method that defines the boolean value of an object -18. `__format__`: Method that returns a formatted string representation of an object - -## 106- List some of the dunder methods used in Mathematical operations - -- `__add__`: Method that defines the behavior of the + operator -- `__sub__`: Method that defines the behavior of the - operator -- `__mul__`: Method that defines the behavior of the * operator -- `__truediv__`: Method that defines the behavior of the / operator (`__div__` was the Python 2 name; it is never called in Python 3) -- `__floordiv__`: Method that defines the behavior of the // operator -- `__mod__`: Method that defines the behavior of the % operator -- `__pow__`: Method that defines the behavior of the ** operator -- `__matmul__`: Method that defines the behavior of the @ operator (matrix multiplication) -- `__neg__`: Method that defines the behavior of unary minus (`-x`) - -Each binary operator also has a reflected form (`__radd__`, `__rsub__`, …), tried when the left operand doesn't know how to handle the right one, and an in-place form (`__iadd__` for `+=`, …). +1. `__init__`: Initializes a newly created instance (the object itself is created by `__new__`, the rarely overridden true constructor) +2. `__repr__`: Returns the unambiguous, developer-facing string representation of an object (used by `repr()`) +3. `__str__`: Returns the readable, user-facing string representation of an object (used by `str()` and `print()`) +4. `__len__`: Returns the length of an object (used by `len()`) +5. `__getitem__`: Gets an item using square brackets (`obj[key]`) +6. `__setitem__`: Sets an item using square brackets (`obj[key] = value`) +7. `__delitem__`: Deletes an item using square brackets (`del obj[key]`) +8. `__iter__`: Returns an iterator over the object (used by `iter()` and `for` loops) +9. `__next__`: Returns the next value from an iterator (used by `next()`) +10. `__call__`: Lets you call an object as if it were a function (`obj()`) +11. `__getattr__`: Called when normal attribute lookup fails to find an attribute +12. `__setattr__`: Called whenever an attribute is assigned (`obj.x = value`) +13. `__delattr__`: Called whenever an attribute is deleted (`del obj.x`) +14. `__enter__`: Called when a `with` block is entered +15. `__exit__`: Called when a `with` block is exited, even if an exception was raised +16. `__hash__`: Returns the hash value of an object (used by `hash()`, dictionaries, and sets) +17. `__bool__`: Defines the truth value of an object (used by `bool()` and `if`) +18. `__format__`: Returns a formatted string representation of an object (used by `format()` and f-strings) + +## 106- List some of the dunder methods used in mathematical operations + +- `__add__`: Defines the `+` operator +- `__sub__`: Defines the `-` operator +- `__mul__`: Defines the `*` operator +- `__truediv__`: Defines the `/` operator (`__div__` was the Python 2 name; it is never called in Python 3) +- `__floordiv__`: Defines the `//` operator +- `__mod__`: Defines the `%` operator +- `__pow__`: Defines the `**` operator +- `__matmul__`: Defines the `@` operator (matrix multiplication) +- `__neg__`: Defines unary minus (`-x`) + +Each binary operator also has a reflected form (`__radd__`, `__rsub__`, …), which Python tries when the left operand doesn't know how to handle the right one, and an in-place form (`__iadd__` for `+=`, …). See question 130 for how this dispatch works. Comparison operators get their own dunders (not strictly "mathematical", but usually asked together): -- `__eq__`: Method that defines the behavior of the == operator -- `__ne__`: Method that defines the behavior of the != operator -- `__lt__`: Method that defines the behavior of the < operator -- `__le__`: Method that defines the behavior of the <= operator -- `__gt__`: Method that defines the behavior of the > operator -- `__ge__`: Method that defines the behavior of the >= operator +- `__eq__`: Defines the `==` operator +- `__ne__`: Defines the `!=` operator +- `__lt__`: Defines the `<` operator +- `__le__`: Defines the `<=` operator +- `__gt__`: Defines the `>` operator +- `__ge__`: Defines the `>=` operator (`functools.total_ordering` fills in the rest if you define `__eq__` plus any one ordering method.) -## 107- List some of the Dunder variables +## 107- List some of the dunder variables - `__name__`: The name of the current module - `__file__`: The path to the file that the module was loaded from @@ -4298,27 +4318,27 @@ Comparison operators get their own dunders (not strictly "mathematical", but usu - `__annotations__`: A dictionary containing type annotations for the module, function, class, or method - `__package__`: The name of the package that the module belongs to - `__loader__`: The loader that loaded the module -- `__spec__`: The specification of the module -- `__path__`: The path to the package that the module belongs to (if it is a package) +- `__spec__`: The module spec (a `ModuleSpec` object), which describes how the module was found and loaded +- `__path__`: For packages only: the list of directories searched for the package's submodules - `__builtins__`: Access to the built-in names (a CPython implementation detail: it is the `builtins` module itself in `__main__`, but a plain dict inside imported modules — `import builtins` is the portable way to reach it) - `__all__`: A list of strings containing the names of the symbols that should be exported when using the `from module import *` syntax -- `__dict__`: A dictionary or other mapping object used to store an object’s attributes. +- `__dict__`: A dictionary (or other mapping) that stores an object's attributes -## 108- What are python frameworks for web development? +## 108- What are the main Python frameworks for web development? -1. **Django**: A high-level, open-source web framework following the MTV (model-template-view) pattern — Django's own naming of what is essentially MVC, with the framework itself acting as controller. It's a batteries-included framework (ORM, migrations, admin interface, auth, forms) known for its robust and scalable approach to web development. +1. **Django**: A high-level, open-source web framework that follows the MTV (model-template-view) pattern — Django's own name for what is essentially MVC, with the framework itself acting as the controller. It is a "batteries-included" framework (ORM, migrations, admin interface, authentication, forms) known for its robust and scalable approach to web development. -2. **Flask**: A lightweight, micro web framework that is easy to use and great for small to medium-sized web applications. It is easy to learn and start with and gives developers more control over the application's structure and behavior. +2. **Flask**: A lightweight "micro" framework that provides routing and request handling and leaves everything else (database access, forms, authentication) to extensions you choose. It is easy to learn, works well for small to medium-sized applications, and gives developers more control over the application's structure. 3. **Pyramid**: A web framework designed for small and large web applications. It provides a lot of flexibility and can be used for a wide range of applications, from small personal blogs to large enterprise applications. -4. **Tornado**: A web framework for building web applications that are highly concurrent and perform well under heavy loads. It is ideal for building real-time applications, such as web sockets and long polling applications. +4. **Tornado**: A web framework and asynchronous networking library for applications that handle many simultaneous connections under heavy load. It is well suited to real-time applications, such as WebSocket and long-polling applications. -5. **FastAPI**: A modern, fast, web framework for building APIs with Python 3.6+ based on standard Python type hints. FastAPI is built on top of Starlette for web parts and Pydantic for data parts. +5. **FastAPI**: A modern, high-performance framework for building APIs, based on standard Python type hints. It is built on top of Starlette (the web layer) and Pydantic (data validation), and it generates OpenAPI documentation automatically (see question 149). -6. **Flask-RESTful**: A simple but flexible extension for Flask that makes it easy to handle RESTful API requests. +6. **Flask-RESTful**: A simple but flexible extension for Flask that makes it easy to build REST APIs. -7. **Sanic**: A Flask like Python 3.5+ web server that's written to go fast. It allows the usage of the async/await syntax added in Python 3.5, which makes your code non-blocking and speedy. +7. **Sanic**: A Flask-like asynchronous framework and web server built for speed. It uses the `async`/`await` syntax, so request handlers don't block while they wait for I/O. ## 109- Write an API using Django REST @@ -4363,18 +4383,17 @@ urlpatterns = [ ] ``` -This example defines a simple `Book` model with fields for (`title, author, published date, and price`). The `BookSerializer` class is used to convert the `Book` model into a format that can be returned by the API. The `BookViewSet` class is a view that handles the logic for creating, reading, updating, and deleting books. The URL routing is handled by the `DefaultRouter` class, which automatically generates the appropriate URLs for the API based on the views and models. +This example defines a simple `Book` model with `title`, `author`, `published_date`, and `price` fields. The `BookSerializer` class converts `Book` instances to and from JSON and validates incoming data. The `BookViewSet` class is a view that handles creating, reading, updating, and deleting books. URL routing is handled by the `DefaultRouter` class, which automatically generates the URLs for every action of the viewset. -Wiring it up requires adding `'rest_framework'` and `'myapp'` to `INSTALLED_APPS`, including `myapp.urls` from the project's root `urls.py`, and running `python manage.py makemigrations && python manage.py migrate` to create the table. +To wire it up, add `'rest_framework'` and `'myapp'` to `INSTALLED_APPS`, include `myapp.urls` from the project's root `urls.py`, and run `python manage.py makemigrations && python manage.py migrate` to create the database table. -You can now run the development server and access the API endpoints at `http://localhost:8000/books/`. Because `ModelViewSet` bundles all the CRUD actions, the router exposes the full REST surface: `GET /books/` (list), `POST /books/` (create), and `GET/PUT/PATCH/DELETE /books//` (retrieve, update, partial update, delete) — plus the browsable HTML API for free during development. +You can now run the development server and access the API endpoints at `http://localhost:8000/books/`. Because `ModelViewSet` bundles all the CRUD (create, read, update, delete) actions, the router exposes all the standard REST endpoints: `GET /books/` (list), `POST /books/` (create), and `GET/PUT/PATCH/DELETE /books//` (retrieve, update, partial update, delete) — plus the browsable HTML API for free during development. -## 110- What are the differences between Django Framework and Django REST Framework? +## 110- What are the differences between Django and Django REST Framework? -Django Framework and Django REST Framework are both web frameworks built on top of the Python programming language. However, they have different purposes and use different approaches to building web applications. **Django Framework** is a general-purpose web framework that can be used to build any type of web application. -**Django REST Framework** is a framework specifically designed for building RESTful APIs. It provides a number of features that make it easy to build APIs that follow the REST architectural style. +They have different purposes. **Django** is a general-purpose web framework that can be used to build any type of web application, typically one that renders HTML pages on the server. **Django REST Framework (DRF)** is not a separate framework: it is a toolkit that runs on top of Django and adds what you need to build RESTful APIs that follow the REST architectural style. -In practice DRF adds, on top of plain Django: **serializers** (declarative conversion and validation between models and JSON), **generic views/viewsets + routers** (CRUD endpoints in a few lines), **authentication schemes** (token, session, and easy JWT integration), fine-grained **permissions and throttling**, **pagination and filtering**, content negotiation, and the **browsable API** — an HTML interface for exploring endpoints during development. +In practice, DRF adds, on top of plain Django: **serializers** (declarative conversion and validation between models and JSON), **generic views/viewsets + routers** (CRUD endpoints in a few lines), **authentication schemes** (token, session, and easy JWT integration), fine-grained **permissions and throttling**, **pagination and filtering**, content negotiation, and the **browsable API** — an HTML interface for exploring endpoints during development. The contrast shows in code. Hand-rolling a JSON endpoint in plain Django (note how serialization, and eventually validation, pagination, and auth, are all yours to write): @@ -4393,7 +4412,7 @@ class SerializedListView(View): return HttpResponse(json_data, content_type='application/json') ``` -The equivalent in Django REST Framework — list **and** create, with validation, auth and pagination hooked in by configuration: +The equivalent in Django REST Framework — list **and** create, with validation, authentication, and pagination enabled through configuration: ```python from rest_framework import generics, permissions @@ -4408,9 +4427,11 @@ class MyObjListCreateAPIView(generics.ListCreateAPIView): serializer_class = MyObjSerializer ``` -The rule of thumb: server-rendered HTML sites need only Django; the moment the deliverable is a JSON API consumed by a SPA or mobile app, DRF (or a framework like FastAPI) earns its place. +The rule of thumb: server-rendered HTML sites need only Django; as soon as the deliverable is a JSON API consumed by a single-page application (SPA) or a mobile app, DRF (or a framework like FastAPI) is worth adding. + +## 111- Create an LRU cache using the `OrderedDict` class -## 111- Create a `LRU Caching` using OrderedDict class +An LRU (least recently used) cache holds a fixed number of items; when it is full, it evicts the item that was used least recently. `OrderedDict` makes this easy because it keeps keys in order and can move a key to the end in O(1) time: ```python from collections import OrderedDict @@ -4450,27 +4471,30 @@ print(lru_cache.get(3)) # Returns 3 print(lru_cache.get(4)) # Returns 4 ``` -Every operation is O(1): the dict lookup, `move_to_end` (a doubly-linked-list splice — the exact capability `OrderedDict` retains over a plain dict), and `popitem(last=False)` for evicting the least-recently-used entry at the front. This is the standard interview implementation; when you just need function-result caching rather than an explicit cache object, `functools.lru_cache` gives you the same policy as a decorator. +Every operation is O(1): the dict lookup, `move_to_end` (which relinks a node in `OrderedDict`'s internal doubly linked list — the capability it still has over a plain dict), and `popitem(last=False)` for evicting the least recently used entry at the front. This is the standard interview implementation; when you just need function-result caching rather than an explicit cache object, `functools.lru_cache` gives you the same policy as a decorator. + +## 112- In a peaceful kingdom, there are houses numbered from 1 to n. The king announces a prize of 100 gold coins for special groups of houses. Your task is to determine how many groups of three houses are special, where a group is special if the sum of the squares of the two smaller house numbers equals the square of the largest house number -## 112- In a peaceful kingdom, there are houses numbered from 1 to n. The king announces a prize of 100 gold coins to some special group of houses. You have been given the task to determine how many sets of three houses can form a special group, where the sum of the squares of two smaller house numbers is equal to the square of the largest house number +In other words, count the Pythagorean triples (a² + b² = c²) whose members are all at most `n`. ```python -#Example 1: -#Input: n = 10 -#Output: 4 -#Explanation: Among the houses numbered 1 to 10, four groups of houses (3, 4, 5), (4, 3, 5), (6, 8, 10), and (8, 6, 10) form special group where sum of square of two smaller houses is equal to the square of larger house. +# Example 1: +# Input: n = 10 +# Output: 4 +# Explanation: Among the houses numbered 1 to 10, four groups of houses, (3, 4, 5), (4, 3, 5), (6, 8, 10), and (8, 6, 10), +# form a special group, where the sum of the squares of the two smaller numbers equals the square of the largest. -#Example 2: -#Input: n = 5 -#Output: 2 -#Explanation: (3,4,5) and (4,3,5). +# Example 2: +# Input: n = 5 +# Output: 2 +# Explanation: (3, 4, 5) and (4, 3, 5). def count_special_groups(n): - squares = {i*i: True for i in range(1, n+1)} + squares = {i * i for i in range(1, n + 1)} count = 0 - for c in range(1, n+1): + for c in range(1, n + 1): c_sq = c * c for a in range(1, c): a_sq = a * a @@ -4510,50 +4534,52 @@ print(result) # Output: 4 # Time Complexity: O(n^2) | Space Complexity: O(1) ``` -Note that each Pythagorean triple is counted twice — `(3, 4, 5)` and `(4, 3, 5)` are distinct ordered groups per the problem statement's own examples. Counting unordered triples instead just means iterating `a` only up to `math.isqrt(c_sq // 2)` (i.e. requiring `a < b`), which for `n = 10` gives `2`: `(3, 4, 5)` and `(6, 8, 10)`. Use `math.isqrt` rather than `int(math.sqrt(...))` for exactness — floating-point `sqrt` misclassifies large perfect squares. +Note that each Pythagorean triple is counted twice — `(3, 4, 5)` and `(4, 3, 5)` are distinct ordered groups per the problem statement's own examples. Counting unordered triples instead just means iterating `a` only up to `math.isqrt(c_sq // 2)` (i.e., requiring `a < b`), which for `n = 10` gives `2`: `(3, 4, 5)` and `(6, 8, 10)`. Use `math.isqrt` rather than `int(math.sqrt(...))` for exact results — floating-point `sqrt` can misjudge whether a large number is a perfect square. ## 113- Write a solution for the Max Pairwise Product Problem +The problem: given a list of numbers, find the largest product of two elements at different positions in the list. + ```python def max_pairwise_product(arr): if len(arr) < 2: raise ValueError("Array must have at least two elements.") - - # Track two largest positive and two smallest negative numbers + + # Track the two largest and the two smallest numbers max1 = max2 = float('-inf') min1 = min2 = float('inf') - + for num in arr: if num > max1: max2 = max1 max1 = num elif num > max2: max2 = num - + if num < min1: min2 = min1 min1 = num elif num < min2: min2 = num - + return max(max1 * max2, min1 * min2) # Time Complexity: O(n) | Space Complexity: O(1) ``` -This code handles primary case and edge cases like: +This code handles the general case and edge cases such as: -- Array with less than two elements. -- Array with negative numbers (e.g., −10, −20, −30 → product of two smallest negatives is the largest positive). -- Array with duplicates (e.g., 5, 5, 5 → product of two largest is 25). +- An array with fewer than two elements. +- An array with negative numbers (e.g., −10, −20, −30 → the product of the two smallest negatives is the largest positive result). +- An array with duplicates (e.g., 5, 5, 5 → the product of the two largest is 25). -The reasoning behind `max(max1 * max2, min1 * min2)`: the maximum pairwise product comes either from the two largest values or — when large-magnitude negatives exist — from the two smallest, whose product is positive. A single O(n) pass tracking those four values beats the two obvious alternatives: the brute-force double loop, O(n²), and sorting first, O(n log n) — though the sorted version `max(a[0] * a[1], a[-1] * a[-2])` (or `heapq.nlargest(2, ...)`/`nsmallest(2, ...)`) is a fine first answer before optimising. In the classic statement of this problem the inputs are non-negative, where the two largest alone suffice; the negative-number handling here generalises it. +The reasoning behind `max(max1 * max2, min1 * min2)`: the maximum pairwise product comes either from the two largest values or — when large-magnitude negatives exist — from the two smallest, whose product is positive. A single O(n) pass tracking those four values beats the two obvious alternatives: the brute-force double loop, O(n²), and sorting first, O(n log n) — though the sorted version `max(a[0] * a[1], a[-1] * a[-2])` (or `heapq.nlargest(2, ...)`/`nsmallest(2, ...)`) is a fine first answer before optimizing. In the classic statement of this problem, the inputs are non-negative, so the two largest values alone suffice; the handling of negative numbers here generalizes it. ## 114- What is the difference between `__new__` and `__init__`? And how does object construction work? -They are two distinct steps of a two-phase construction protocol, and conflating them is a common source of confusion. +They are two separate steps of object construction, and mixing them up is a common source of confusion. - **`__new__(cls, ...)` is the actual constructor**: a static method that **allocates and returns** the new instance. It runs _first_, and its job is to produce the object. -- **`__init__(self, ...)` is the initialiser**: it receives the _already-created_ instance as `self`, configures its attributes, and **returns `None`**. It never creates anything. +- **`__init__(self, ...)` is the initializer**: it receives the _already-created_ instance as `self`, configures its attributes, and **returns `None`**. It never creates anything. When you call `MyClass(args)`, it is the metaclass's `type.__call__` that orchestrates the sequence: it calls `__new__` to get the instance, and then — **only if `__new__` returned an instance of `cls`** — calls `__init__` on it. @@ -4565,12 +4591,12 @@ class Demo: return instance def __init__(self, value): - print("2. __init__ - initialising the instance") + print("2. __init__ - initializing the instance") self.value = value d = Demo(42) # 1. __new__ - creating the instance -# 2. __init__ - initialising the instance +# 2. __init__ - initializing the instance ``` The subtle rule to state explicitly: **if `__new__` returns an object that is not an instance of `cls`, `__init__` is skipped entirely.** This is the mechanism behind returning a cached or different object. @@ -4589,7 +4615,7 @@ You rarely override `__new__` — `__init__` covers almost every case. The legit print(PositiveInt(5) + 10) # 15 - behaves as an int ``` -- **Singletons / instance caching / object pools** — return an existing instance instead of a fresh one (note the caveat that `__init__` still re-runs on the returned instance if it _is_ a `cls` instance, so guard against re-initialising): +- **Singletons / instance caching / object pools** — return an existing instance instead of a fresh one (note the caveat that `__init__` still re-runs on the returned instance if it _is_ a `cls` instance, so guard against re-initializing): ```python class Singleton: @@ -4602,7 +4628,7 @@ You rarely override `__new__` — `__init__` covers almost every case. The legit print(Singleton() is Singleton()) # True ``` -- **Metaclasses** (`type.__new__` customises class creation), and factory patterns that return a subclass chosen at runtime. +- **Metaclasses** (`type.__new__` customizes class creation), and factory patterns that return a subclass chosen at runtime. Two footnotes worth mentioning: `dataclasses`, `NamedTuple`, and most everyday classes never touch `__new__`; and `__new__` is implicitly a static method even though you don't decorate it, which is why it takes `cls` explicitly rather than receiving it like a classmethod. @@ -4610,9 +4636,9 @@ Two footnotes worth mentioning: `dataclasses`, `NamedTuple`, and most everyday c The unifying idea first: **all three store _references_ (pointers) to objects, never the objects inline.** `sys.getsizeof(container)` therefore measures the container's own bookkeeping and pointer array, not the elements it points at — a list of a million ints and a list of a million huge strings have the _same_ `getsizeof`. -**`list` — a dynamic array with a separately-allocated buffer.** A `PyListObject` is a small fixed header holding: a pointer to a heap-allocated **array of `PyObject*`**, `ob_size` (the number of elements in use), and `allocated` (the array's current capacity). Because the buffer is separate and resizable, a list can grow and shrink; it **over-allocates** spare capacity so appends are amortised O(1) (see the next question). Indexing is O(1) pointer arithmetic; inserting/deleting at the front is O(n) because everything shifts. +**`list` — a dynamic array with a separately allocated buffer.** A `PyListObject` is a small fixed header holding: a pointer to a heap-allocated **array of `PyObject*`**, `ob_size` (the number of elements in use), and `allocated` (the array's current capacity). Because the buffer is separate and resizable, a list can grow and shrink; it **over-allocates** spare capacity so appends are amortized O(1) (see the next question). Indexing is O(1) pointer arithmetic; inserting/deleting at the front is O(n) because everything shifts. -**`tuple` — a fixed array stored inline.** A `PyTupleObject` stores its pointer array **in the same single allocation** as the header (a variable-length object), with no separate buffer and, crucially, **no `allocated` slack** — an immutable tuple is sized exactly once at creation and never grows. That makes a tuple **smaller and slightly faster to build and access** than an equivalent list. CPython goes further: it keeps **free lists** of small tuples for fast reuse, and tuples of constants are cached/interned in compiled code. +**`tuple` — a fixed array stored inline.** A `PyTupleObject` stores its pointer array **in the same single allocation** as the header (a variable-length object), with no separate buffer and, crucially, **no `allocated` slack** — an immutable tuple is sized exactly once at creation and never grows. That makes a tuple **smaller and slightly faster to build and access** than an equivalent list. CPython goes further: it keeps **free lists** of recently freed small tuples for fast reuse, and tuples made only of constants are built once at compile time and stored with the code. ```python import sys @@ -4622,11 +4648,11 @@ print(sys.getsizeof((1, 2, 3))) # e.g. 64 - header + inline pointers, no sla **`set` — a hash table without values.** A `set` is essentially a `dict` that stores only keys: an open-addressing hash table of slots, each holding a reference and the element's cached hash. It resizes when it passes a load-factor threshold (roughly 3/5 full), trading memory for O(1) membership. Consequences: elements must be **hashable**, there is **no order** guarantee, and a set uses **more** memory than a list of the same elements because most slots sit empty to keep collisions rare. That empty space is exactly what buys O(1) `in` versus a list's O(n). -The practical takeaway for a senior engineer: reach for a **tuple** for fixed, immutable records (smaller, hashable, usable as dict keys), a **list** when you need to grow/mutate an ordered sequence, and a **set** the moment you do repeated membership tests or need de-duplication. For large homogeneous numeric data, none of these are ideal — `array.array` (inline C values, no per-element object) or NumPy (contiguous typed buffer) avoid the per-element pointer-and-object overhead entirely. +The practical takeaway for a senior engineer: use a **tuple** for fixed, immutable records (smaller, hashable, usable as dict keys), a **list** when you need to grow or modify an ordered sequence, and a **set** the moment you do repeated membership tests or need de-duplication. For large homogeneous numeric data, none of these are ideal — `array.array` (inline C values, no per-element object) or NumPy (contiguous typed buffer) avoid the per-element pointer-and-object overhead entirely. ## 116- When a list grows beyond its allocated capacity, what happens to the underlying array allocation? And how are the existing elements handled? -This is the mechanism that makes `list.append` **amortised O(1)**, and it is worth being precise about. +This is the mechanism that makes `list.append` **amortized O(1)**, and it is worth being precise about. A list keeps two numbers: `ob_size` (elements currently used) and `allocated` (slots the backing buffer can hold). As long as `ob_size < allocated`, an append just writes into the next free slot — genuinely O(1). The interesting case is when `ob_size == allocated` and you append again: @@ -4636,7 +4662,7 @@ A list keeps two numbers: `ob_size` (elements currently used) and `allocated` (s new_allocated = new_size + (new_size >> 3) + 6 # then rounded ``` - i.e. about **12.5% headroom** on top of what's needed, plus a small constant, with the result rounded to a multiple of 4. The resulting capacity progression is 0, 4, 8, 16, 24, 32, 40, 52, 64, 76, 92, 108, … — geometric-ish growth, not linear. + i.e., about **12.5% headroom** on top of what's needed, plus a small constant, with the result rounded to a multiple of 4. The resulting capacity progression is 0, 4, 8, 16, 24, 32, 40, 52, 64, 76, 92, 108, … — roughly geometric growth, not linear. 2. It calls **`realloc`** on the pointer buffer to that new capacity. Two things can happen under the hood: the allocator may **extend the block in place** (no copy needed), or, if the adjacent memory is taken, it **allocates a fresh, larger block, copies the existing pointers over, and frees the old block.** @@ -4654,14 +4680,14 @@ for i in range(20): lst.append(i) ``` -**Why this gives amortised O(1):** because capacity grows geometrically, the total cost of the occasional O(n) copies across `n` appends sums to O(n), so the _average_ cost per append is O(1). Any single append that triggers a resize is O(n), but those are rare and get rarer as the list grows — the classic amortised-analysis result. +**Why this gives amortized O(1):** because capacity grows geometrically, the total cost of the occasional O(n) copies across `n` appends sums to O(n), so the _average_ cost per append is O(1). Any single append that triggers a resize is O(n), but those are rare and get rarer as the list grows — the classic amortized-analysis result. Senior-level corollaries: - **Identity is preserved for elements, not for the buffer.** After a resize, `lst[0] is original_first_element` is still `True`, but the internal buffer address changed — which is why holding a raw pointer into a list from a C extension across an append is a bug. -- **Pre-size when you know the length.** Building `[None] * n` and assigning by index, or using a list comprehension, avoids repeated regrowth. `list.extend(iterable)` also grows once (using the iterable's length hint) rather than reallocating per element. +- **Pre-size when you know the length.** Building `[None] * n` and assigning by index avoids repeated regrowth. `list(iterable)` and `list.extend(iterable)` also grow the buffer once, up front, when the iterable can report its length, rather than reallocating as elements arrive. - **Lists also shrink**: deletions that drop usage well below capacity trigger a realloc to a smaller buffer, so a list that was huge and is now small releases most of its buffer. -- **This is why front operations are the wrong tool.** `insert(0, x)` and `pop(0)` are O(n) regardless of capacity because they shift every pointer. For a queue, use `collections.deque`, which is a linked list of fixed-size blocks with O(1) appends/pops at both ends and never does this whole-array copy. +- **This is why front operations are the wrong tool.** `insert(0, x)` and `pop(0)` are O(n) regardless of capacity because they shift every pointer. For a queue, use `collections.deque`, which is a doubly linked list of fixed-size blocks with O(1) appends and pops at both ends and never does this whole-array copy. ## 117- What hashing is used for `dict` keys, and how do `__hash__` and `__eq__` interact? @@ -4670,7 +4696,7 @@ A `dict` (and `set`) is a **hash table**, so every key must be **hashable**: it **How a lookup actually works.** For `d[key]`, CPython: 1. Computes `hash(key)` and uses the low bits to pick a slot in the table. -2. If the slot is occupied, it compares the stored key to `key` — first with an `is` **identity** short-circuit (fast path: the same object is trivially equal), then with `==`. This comparison is what resolves **collisions**: two different keys can land in the same slot, and `__eq__` decides whether you found your key or a colliding neighbour. +2. If the slot is occupied, it compares the stored key to `key` — first with an `is` **identity** short-circuit (fast path: the same object is trivially equal), then with `==`. This comparison is what resolves **collisions**: two different keys can land in the same slot, and `__eq__` decides whether you found your key or a different key that happens to share the slot. 3. On a collision-and-mismatch, it **probes** further slots (CPython perturbs the probe sequence with more hash bits, an open-addressing scheme) until it finds the key or an empty slot. So both methods are used on every lookup: `__hash__` to _locate_ the bucket, `__eq__` to _confirm_ the match. @@ -4702,12 +4728,12 @@ print(pts[Point(1, 2)]) # 'a' - a different object, equal value, found correct **Two facts a senior is expected to know:** -- **String and bytes hashing is randomised per process** (SipHash, since Python 3.3, controlled by `PYTHONHASHSEED`). This defends against algorithmic-complexity (hash-flooding) DoS attacks, where an attacker sends keys engineered to collide and degrade a dict to O(n). It also means `hash("x")` differs between runs — never persist or depend on raw hash values across processes. (Dict _insertion order_ is still preserved regardless; that's unrelated to hashing.) +- **String and bytes hashing is randomized per process** (on by default since Python 3.3, using the SipHash algorithm since 3.4, and controlled by the `PYTHONHASHSEED` environment variable). This defends against hash-flooding denial-of-service attacks, in which an attacker sends keys crafted to collide, degrading every dict lookup to O(n). It also means `hash("x")` differs between runs — never persist or depend on raw hash values across processes. (Dict _insertion order_ is still preserved regardless; that's unrelated to hashing.) - **Average lookup is O(1); worst case is O(n)** if hashing degenerates and everything collides. Small integers hash to themselves (`hash(5) == 5`), and `hash(-1)` is special-cased to `-2` because `-1` signals an error in the C API. ## 118- What is `asyncio`, and how do `async`/`await` and the event loop actually work? -`asyncio` is Python's framework for **single-threaded concurrency via cooperative multitasking**. The whole model turns on one distinction: it gives you concurrency _without_ threads by letting one thread juggle thousands of tasks, switching between them only at explicit `await` points. +`asyncio` is Python's framework for **single-threaded concurrency via cooperative multitasking**: tasks voluntarily hand control back instead of being interrupted by the operating system. It gives you concurrency _without_ extra threads by letting one thread juggle thousands of tasks, switching between them only at explicit `await` points. - An **`async def`** function is a **coroutine function**; calling it does _not_ run it — it returns a **coroutine object**, much as calling a generator function returns a generator. Nothing executes until the coroutine is driven by an event loop. - **`await`** suspends the current coroutine until the awaited awaitable completes, and — this is the key part — **hands control back to the event loop** so it can run other ready tasks meanwhile. `await` is a cooperative yield point, not a blocking wait. @@ -4732,12 +4758,12 @@ asyncio.run(main()) # ~3s total, not 6 - the sleeps overlap The three `sleep`s overlap because each `await asyncio.sleep` yields control, so the loop starts the next task instead of waiting. Total time is the _longest_ task, not the sum — that's the payoff. -**The non-negotiable rule:** the event loop runs on one thread, so **a blocking call blocks everything.** A plain `time.sleep(2)`, a synchronous `requests.get`, or a CPU-heavy loop inside a coroutine freezes _all_ tasks, because nothing yields back to the loop. You must use async-aware equivalents (`asyncio.sleep`, `aiohttp`/`httpx`, async database drivers), or offload blocking work with `await asyncio.to_thread(func)` (thread pool) or a process pool for CPU-bound work. +**The critical rule:** the event loop runs on one thread, so **a blocking call blocks everything.** A plain `time.sleep(2)`, a synchronous `requests.get`, or a CPU-heavy loop inside a coroutine freezes _all_ tasks, because nothing yields back to the loop. You must use async-aware equivalents (`asyncio.sleep`, `aiohttp`/`httpx`, async database drivers), or offload blocking work with `await asyncio.to_thread(func)` (thread pool) or a process pool for CPU-bound work. Where it fits: -- **Ideal for high-concurrency I/O-bound workloads** — thousands of simultaneous network connections, API calls, or socket clients — where threads would waste memory (each thread has a full stack) and the GIL makes their parallelism moot anyway. One event-loop thread handling 10,000 sockets is the canonical win. -- **Useless for CPU-bound work.** Coroutines don't sidestep the GIL and there's only one thread; heavy computation needs `multiprocessing`. +- **Ideal for high-concurrency I/O-bound workloads** — thousands of simultaneous network connections, API calls, or socket clients — where threads would waste memory (each thread has its own full stack) and the GIL prevents them from running Python code in parallel anyway. One event-loop thread handling 10,000 sockets is the classic success story. +- **Useless for CPU-bound work.** Coroutines don't avoid the GIL, and there's only one thread; heavy computation needs `multiprocessing`. - **`await` composes; `create_task` fans out.** `await coro` runs sequentially; `asyncio.create_task(coro)` schedules it to run concurrently and returns a `Task` you can await later; `asyncio.gather`/`asyncio.TaskGroup` (3.11+) run many concurrently and collect results. Mental model to close on: `asyncio` is **not parallelism** — it's one worker interleaving many jobs by never sitting idle during I/O. It trades the OS scheduler's preemptive thread switching for explicit, cheap, cooperative switching at `await`, which is why it scales to far more concurrent I/O operations than threads while sidestepping the data races that preemptive threading invites. @@ -4748,15 +4774,15 @@ The decision reduces to one question first — **is the work I/O-bound or CPU-bo | Model | Parallelism | Best for | Cost | | --- | --- | --- | --- | -| **`threading`** | No (GIL serialises bytecode) | I/O-bound, moderate concurrency, blocking libraries | preemptive → needs locks; ~MBs per thread stack | -| **`multiprocessing`** | **Yes** (separate interpreters) | **CPU-bound** work | Process overhead; data must be pickled/IPC'd | +| **`threading`** | No (GIL runs one thread's bytecode at a time) | I/O-bound, moderate concurrency, blocking libraries | Preemptive → needs locks; megabytes of stack per thread | +| **`multiprocessing`** | **Yes** (separate interpreters) | **CPU-bound** work | Process overhead; data must be pickled and sent between processes | | **`asyncio`** | No (one thread) | I/O-bound, **very high** concurrency | Needs async-native libraries; one blocking call stalls all | -**CPU-bound work → `multiprocessing` (or a native extension).** Number crunching, image processing, and parsing gain nothing from threads because the GIL lets only one thread run Python bytecode at a time. Separate _processes_ each have their own interpreter and GIL, so they genuinely run in parallel on multiple cores. The price is that they don't share memory — arguments and results are **pickled** and shipped over IPC — so it pays off only when the compute per task dwarfs that transfer cost. (The alternative is to keep threads but do the heavy lifting in a C extension that releases the GIL, which is exactly what NumPy does.) +**CPU-bound work → `multiprocessing` (or a native extension).** Number crunching, image processing, and parsing gain nothing from threads because the GIL lets only one thread run Python bytecode at a time. Separate _processes_ each have their own interpreter and GIL, so they genuinely run in parallel on multiple cores. The price is that they don't share memory — arguments and results are **pickled** and sent between processes (inter-process communication, or IPC) — so it pays off only when each task's computation takes far longer than that transfer. (The alternative is to keep threads but do the heavy lifting in a C extension that releases the GIL, which is exactly what NumPy does.) **I/O-bound work → `threading` or `asyncio`.** While a thread waits on a socket, disk, or subprocess, it releases the GIL, so other threads make progress. Both models overlap I/O effectively; the choice between them is about scale and ecosystem: -- **`threading`** is the pragmatic choice for **moderate** concurrency (tens to low hundreds of tasks) and, decisively, when you must use **blocking libraries** (a synchronous DB driver, `requests`, legacy SDKs). Its downside is preemptive switching, which can interrupt between any two bytecodes, so shared mutable state needs `Lock`/`Queue` and invites race conditions. +- **`threading`** is the pragmatic choice for **moderate** concurrency (tens to low hundreds of tasks) and, decisively, when you must use **blocking libraries** (a synchronous DB driver, `requests`, legacy SDKs). Its downside is preemptive switching: the interpreter can switch threads between any two bytecode instructions, so shared mutable state needs a `Lock` or a `Queue`, and race conditions are easy to introduce. - **`asyncio`** wins at **massive** concurrency (thousands of simultaneous connections) because tasks are cheap (no per-thread stack) and switching is explicit at `await`. But it demands an **async-native stack top to bottom** (`aiohttp`/`httpx`, async DB drivers); a single blocking call anywhere stalls the whole event loop. The high-level interface for the first two is **`concurrent.futures`**, whose `ThreadPoolExecutor` and `ProcessPoolExecutor` share an identical API — so you can write pool code once and switch models by changing one class name: @@ -4773,7 +4799,7 @@ with ProcessPoolExecutor() as pool: results = list(pool.map(crunch, big_inputs)) ``` -A senior would add: these are not mutually exclusive — a real service often combines them (an `asyncio` web layer that offloads CPU work to a `ProcessPoolExecutor` via `loop.run_in_executor`), and the honest first step is always to **measure** whether the bottleneck is I/O wait or CPU before picking a model. And on the horizon, the free-threaded (no-GIL) builds from PEP 703 will eventually let `threading` deliver true CPU parallelism, changing this calculus. +A senior would add: these are not mutually exclusive — a real service often combines them (an `asyncio` web layer that offloads CPU work to a `ProcessPoolExecutor` via `loop.run_in_executor`), and the honest first step is always to **measure** whether the bottleneck is I/O wait or CPU before picking a model. And the free-threaded (no-GIL) builds from PEP 703 (see question 151) already let `threading` deliver true CPU parallelism on opt-in builds; as the ecosystem adopts them, they will change this calculation. ## 120- What is a metaclass, and when would you actually use one? @@ -4786,7 +4812,7 @@ print(type(Foo())) # - Foo() is an instance of Foo print(isinstance(Foo, type)) # True ``` -Because `type` is callable, you can even create classes dynamically without the `class` statement — `type(name, bases, namespace)` is what the `class` block desugars to: +Because `type` is callable, you can even create classes dynamically without the `class` statement — calling `type(name, bases, namespace)` is essentially what a `class` block does behind the scenes: ```python Dog = type("Dog", (), {"sound": "woof", "speak": lambda self: self.sound}) @@ -4810,7 +4836,7 @@ class Point(metaclass=AutoRepr): print(Point(1, 2)) # Point({'x': 1, 'y': 2}) ``` -**When you'd actually use one — and when you wouldn't.** Metaclasses solve exactly one problem: **customising class creation itself** (validating or registering classes, injecting methods, enforcing conventions across a whole class hierarchy). The canonical real-world users are frameworks: Django's and SQLAlchemy's ORMs use metaclasses to turn declarative `class Model` bodies into database-mapped objects, ABCs use `ABCMeta`, and enums use `EnumMeta`. +**When you'd actually use one — and when you wouldn't.** Metaclasses solve exactly one problem: **customizing class creation itself** (validating or registering classes, injecting methods, enforcing conventions across a whole class hierarchy). The typical real-world users are frameworks: Django's and SQLAlchemy's ORMs use metaclasses to turn declarative `class Model` bodies into database-mapped classes, ABCs use `ABCMeta`, and enums use `EnumType` (formerly `EnumMeta`). But for application code the honest senior answer is **"almost never — reach for something simpler first"**: @@ -4818,11 +4844,11 @@ But for application code the honest senior answer is **"almost never — reach f - **`__set_name__`** covers descriptor-naming needs. - **Class decorators** can rewrite a class after creation and are far easier to read and compose than a metaclass. -The famous Tim Peters line captures the judgement expected of a senior: _"If you wonder whether you need metaclasses, you don't."_ Know what they are and how the `type`/instance chain works — the mechanism underpins ORMs and is a favourite interview probe — but treat writing one in ordinary code as a red flag rather than a flex. One practical gotcha to mention: metaclass conflicts. If two base classes have different (non-subclass-related) metaclasses, Python refuses to create the derived class. +The famous Tim Peters line captures the judgment expected of a senior: _"If you wonder whether you need metaclasses, you don't."_ Know what they are and how the `type`/instance chain works — the mechanism underpins ORMs and is a favorite interview topic — but treat writing one in ordinary code as a red flag rather than something to show off. One practical gotcha to mention: metaclass conflicts. If two base classes have unrelated metaclasses (neither is a subclass of the other), Python refuses to create the derived class. ## 121- Do Python's type hints do anything at runtime? -The blunt answer that separates people who've _used_ type hints from people who've only read about them: **by default, no — the interpreter does not enforce them at all.** Annotations are hints for humans and _external_ tools; passing the "wrong" type runs happily until something unrelated breaks. +The short answer: **by default, no — the interpreter does not enforce them at all.** Annotations are hints for humans and _external_ tools; code that passes the "wrong" type runs anyway, until something fails later for a seemingly unrelated reason. ```python def add(a: int, b: int) -> int: @@ -4846,18 +4872,18 @@ print(greet.__annotations__) # {'name': , 'return': Details a senior is expected to have hit in practice: -- **Annotations can be strings (lazy evaluation).** `from __future__ import annotations` (PEP 563) makes _all_ annotations strings that aren't evaluated at definition time — which fixes forward references (referring to a class not yet defined) and circular-import issues, but means anything reading them at runtime must use `get_type_hints()` to resolve them. This is a genuine friction point with Pydantic and other introspecting libraries. +- **Annotations can be strings (lazy evaluation).** `from __future__ import annotations` (PEP 563) makes _all_ annotations strings that aren't evaluated at definition time — which fixes forward references (referring to a class not yet defined) and circular-import issues, but means anything reading them at runtime must use `get_type_hints()` to resolve them. This is a genuine source of friction with Pydantic and other libraries that inspect annotations at runtime. Python 3.14 (PEP 649) made all annotations lazily evaluated by default, which solves the forward-reference problem without turning them into strings. - **Generics and the typing toolkit:** `list[int]`, `dict[str, int]`, `Optional[X]` (= `X | None`), `Union` / the `X | Y` syntax (3.10+), `TypeVar` and `Generic` for parametric code, `Protocol` for structural typing, `Callable`, `Any`, `Literal`, `TypedDict`, and `cast`. `typing.TYPE_CHECKING` guards imports needed only for hints. -- **`Any` disables checking** for that value — useful as an escape hatch, dangerous as a habit. +- **`Any` disables checking** for that value — useful as a deliberate opt-out, dangerous as a habit. - **Hints don't affect performance** meaningfully; they're metadata, not runtime checks. -The senior framing: type hints are **gradual and optional** — you add them where they earn their keep (public APIs, complex data flows, large teams) and a static checker in CI turns them into a real safety net. But never assume the interpreter is validating them; if you need runtime guarantees, that's a job for Pydantic, explicit checks, or `assert isinstance(...)`. +The senior framing: type hints are **gradual and optional** — you add them where they pay off (public APIs, complex data flows, large teams), and a static checker in CI turns them into a real safety net. But never assume the interpreter is validating them; if you need runtime guarantees, that's a job for Pydantic, explicit checks, or `assert isinstance(...)`. ## 122- What is the walrus operator (`:=`) and when is it useful? The walrus operator `:=` (named for its resemblance to a walrus's eyes and tusks), introduced in Python 3.8 by PEP 572, performs **assignment inside an expression**. Ordinary `=` is a statement and cannot appear where a value is expected; `:=` both assigns to a name _and_ evaluates to that value, so it can live inside an `if`, `while`, comprehension, or function call. -Its value is eliminating the choice between **computing something twice** and **adding an extra line** — the "assign, then test" and "assign, then use" patterns: +Its value is that you no longer have to choose between **computing something twice** and **adding an extra line** in the "assign, then test" and "assign, then use" patterns: ```python # Without walrus: either call len() twice, or add a setup line before the if @@ -4882,11 +4908,11 @@ The senior perspective is as much about **restraint** as capability: ## 123- What is in `functools` beyond `lru_cache`? -`functools` is the standard library's toolkit for **higher-order functions** — utilities that act on or return other functions. Beyond `lru_cache`/`cache` (covered separately), the pieces a senior reaches for regularly: +`functools` is the standard library's toolkit for **higher-order functions** — utilities that act on or return other functions. Beyond `lru_cache`/`cache` (see question 103), these are the tools a senior engineer uses regularly: -- **`functools.wraps`** — the decorator you apply to a wrapper so it copies the wrapped function's `__name__`, `__doc__`, `__qualname__`, etc. Without it, every decorated function reports itself as `wrapper`, breaking introspection, `help()`, and debuggers. Non-optional in real decorators. +- **`functools.wraps`** — the decorator you apply to a wrapper so it copies the wrapped function's `__name__`, `__doc__`, `__qualname__`, etc. Without it, every decorated function reports itself as `wrapper`, breaking introspection, `help()`, and debuggers. Treat it as mandatory in real decorators. -- **`functools.partial`** — freezes some arguments of a callable, returning a new callable with fewer parameters. The clean way to pre-configure a function for an API that wants a zero/one-arg callback: +- **`functools.partial`** — freezes some arguments of a callable, returning a new callable with fewer parameters. It is the clean way to pre-configure a function for an API that expects a callback with fewer arguments: ```python from functools import partial @@ -4921,7 +4947,7 @@ The senior perspective is as much about **restraint** as capability: - **`functools.cmp_to_key`** — adapts an old-style two-argument comparison function into a `key=` function for `sorted`/`min`/`max`, the bridge from Python 2's `cmp` sorting to Python 3's key-based sorting. -The through-line: `functools` is where Python's functional-programming and metaprogramming conveniences live, and `wraps` + `partial` + `singledispatch` + `cached_property` in particular show up constantly in production library and framework code. +The common thread: `functools` is where Python's functional-programming and metaprogramming conveniences live, and `wraps` + `partial` + `singledispatch` + `cached_property` in particular show up constantly in production library and framework code. ## 124- What are `dataclasses`, and how do they compare to `NamedTuple`, `TypedDict`, and `attrs`? @@ -4943,7 +4969,7 @@ print(p == Point(1, 2)) # True <- generated __eq__ The features worth knowing, because they map straight to real needs: -- **`field(default_factory=...)`** is the sanctioned fix for the mutable-default trap — dataclasses actively **reject** a bare mutable default (`tags: list = []` raises `ValueError`), which is a nice guardrail. +- **`field(default_factory=...)`** is the official fix for the mutable-default trap — dataclasses actively **reject** a bare mutable default (`tags: list = []` raises `ValueError`), which is a helpful safety net. - **`frozen=True`** makes instances immutable and hashable — the correct way to build a value object usable as a dict key or set member. - **`slots=True`** (3.10+) generates `__slots__`, cutting per-instance memory and speeding attribute access for classes instantiated in bulk. - **`order=True`** generates the comparison methods (`<`, `<=`, …) so instances sort by field order. @@ -4955,10 +4981,10 @@ The features worth knowing, because they map straight to real needs: | --- | --- | --- | --- | | **`@dataclass`** | Yes (unless `frozen`) | a normal class instance | general-purpose data holder with methods, defaults, validation | | **`NamedTuple`** | **No** (immutable tuple) | a `tuple` subclass | a lightweight immutable record that should also behave like a tuple (unpackable, indexable) | -| **`TypedDict`** | Yes (it _is_ a dict) | a plain `dict` at runtime | annotating the **shape of a dict** (e.g. a JSON payload) for the type checker, with zero runtime class overhead | +| **`TypedDict`** | Yes (it _is_ a dict) | a plain `dict` at runtime | annotating the **shape of a dict** (e.g., a JSON payload) for the type checker, with zero runtime class overhead | | **`attrs`** | configurable | a normal class instance | you need more power than dataclasses (validators, converters, richer field control) — it's the third-party library dataclasses was inspired by | -Key distinctions to articulate: a **`NamedTuple`** is still a tuple — it's immutable, iterable, and unpackable (`x, y = point`), which dataclasses aren't unless you add it; use it for small fixed records where tuple behaviour is a feature. A **`TypedDict`** creates _no class at all_ at runtime — it's purely a static-typing annotation over an ordinary dict, ideal for typing external JSON without changing how you access it. **`attrs`** predates and outclasses dataclasses in flexibility (field validators, converters, `__slots__` by default), so it's the answer when the stdlib dataclass hits its limits. And when you also need **runtime validation and parsing** of external input, the real production answer is often **Pydantic**, which looks like a dataclass but coerces and validates every field against its annotations. +Key distinctions to articulate: a **`NamedTuple`** is still a tuple — it's immutable, iterable, and unpackable (`x, y = point`), which dataclasses aren't unless you add it; use it for small fixed records where tuple behavior is a feature. A **`TypedDict`** creates _no class at all_ at runtime — it's purely a static-typing annotation over an ordinary dict, ideal for typing external JSON without changing how you access it. **`attrs`** is older than dataclasses and more flexible (field validators, converters, `__slots__` by default), so it's the answer when the standard library's dataclass hits its limits. And when you also need **runtime validation and parsing** of external input, the real production answer is often **Pydantic**, which looks like a dataclass but coerces and validates every field against its annotations. Rule of thumb: `@dataclass` by default; `NamedTuple` for immutable tuple-like records; `TypedDict` to type a dict's shape; `attrs`/Pydantic when you outgrow the standard library. @@ -4973,20 +4999,20 @@ Importing is not textual inclusion — it **executes a module top to bottom exac 3. **Load and execute.** The loader creates a new empty module object, **inserts it into `sys.modules` _before_ executing it** (crucial — see below), then runs the module body, populating its namespace. Compiled bytecode is cached in `__pycache__` to skip recompilation next time. 4. **Bind the name** in the importing namespace (`import x` binds `x`; `from x import y` binds `y`). -**Why circular imports break — and why they sometimes don't.** Because a module is registered in `sys.modules` _before_ its body finishes running, a cycle (A imports B, B imports A) doesn't infinitely recurse — but it can hand you a **half-initialised module**. If A is mid-execution when it triggers B, and B does `from A import thing`, `thing` may not exist yet, giving `ImportError: cannot import name 'thing'` or an `AttributeError`. +**Why circular imports break — and why they sometimes don't.** Because a module is registered in `sys.modules` _before_ its body finishes running, a cycle (A imports B, B imports A) doesn't infinitely recurse — but it can hand you a **half-initialized module**. If A is mid-execution when it triggers B, and B does `from A import thing`, `thing` may not exist yet, giving `ImportError: cannot import name 'thing'` or an `AttributeError`. The senior toolkit for circular imports, in order of preference: - **Restructure** — the cycle usually signals a design problem. Extract the shared piece into a third module both depend on, so the dependency graph becomes acyclic. -- **Import the _module_, not the name.** `import a` and later reference `a.thing` (resolved at call time) instead of `from a import thing` (resolved at import time). This defers the lookup past initialisation. -- **Move the import inside the function** that needs it (a deliberate local import — see the local-vs-global-imports question), so it runs after both modules are fully loaded. +- **Import the _module_, not the name.** `import a` and later reference `a.thing` (resolved at call time) instead of `from a import thing` (resolved at import time). This defers the lookup past initialization. +- **Move the import inside the function** that needs it (a deliberate local import — see question 100), so it runs after both modules are fully loaded. - **For type hints only**, guard the import with `if typing.TYPE_CHECKING:` and use a string annotation — the import never runs at runtime, so it can't cycle. Two related facts worth mentioning: an implicit **namespace package** (a directory without `__init__.py`, PEP 420) is discovered by a different finder and can span multiple `sys.path` entries; and `importlib` is the programmatic API (`importlib.import_module`, `importlib.reload`) when you need to import by name computed at runtime or force a re-execution. ## 126- What is `weakref` and when do you need it? -A **weak reference** points to an object **without incrementing its reference count**, so it does not, by itself, keep the object alive. If the only remaining references to an object are weak, the object is still collected, and the weak references "die" (start returning `None`). This is the escape hatch for the situations where ordinary strong references cause problems. +A **weak reference** points to an object **without incrementing its reference count**, so it does not, by itself, keep the object alive. If the only remaining references to an object are weak, the object is still collected, and the weak references "die" (start returning `None`). This makes weak references the right tool whenever an ordinary (strong) reference would keep an object alive longer than it should. ```python import weakref @@ -5017,13 +5043,13 @@ print(ref()) # None - the object was collected; the weakref is dea return obj ``` -2. **Breaking reference cycles** — parent/child, observer/subject, or doubly-linked back-pointers. If a child holds a strong reference back to its parent, the two form a cycle that reference counting alone can't reclaim (it needs the slower cyclic GC). Making the **back-reference weak** breaks the cycle so plain refcounting collects both promptly, which matters for objects whose cleanup timing you care about. +2. **Breaking reference cycles** — parent/child, observer/subject, or back-references in doubly linked structures. If a child holds a strong reference back to its parent, the two form a cycle that reference counting alone can't reclaim (it needs the slower cyclic GC). Making the **back-reference weak** breaks the cycle so plain refcounting collects both promptly, which matters for objects whose cleanup timing you care about. Details a senior would add: - You **dereference by calling** the weakref (`ref()`), and it returns the object or `None` — always check for `None`, because the object may have vanished between calls. - **`weakref.finalize(obj, callback)`** registers a cleanup to run when the object is collected — a more reliable pattern than `__del__` for "do X when this dies". -- **Not everything is weak-referenceable.** Common built-ins — `int`, `str`, `tuple`, `list`, and `dict` — **cannot** be weakly referenced (though a _subclass_ of them can, because subclassing adds a `__weakref__` slot). A class using `__slots__` also loses weakref support unless it includes `'__weakref__'` in the slots. Ordinary custom classes support it out of the box. +- **Not everything is weak-referenceable.** Common built-ins — `int`, `str`, `tuple`, `list`, and `dict` — **cannot** be weakly referenced. A _subclass_ of `str`, `list`, or `dict` can (subclassing adds a `__weakref__` slot), but in CPython even subclasses of `int` and `tuple` cannot. A class using `__slots__` also loses weakref support unless it includes `'__weakref__'` in the slots. Ordinary custom classes support it out of the box. The mental model: use a weak reference whenever you want to _observe or cache_ an object without _owning_ it — when its lifetime should be decided by someone else, and you want to be notified (via `None`) once it's gone. @@ -5054,7 +5080,7 @@ The kinds of patterns are what make it powerful: - **Class patterns** match by type _and_ pull out attributes positionally or by name: `case Point(x=0, y=y):` matches a `Point` on the y-axis and binds `y`. (Positional matching uses the class's `__match_args__`.) - **Capture, wildcard, OR, and guards:** a bare name captures; `_` matches anything without binding; `|` combines alternatives; and an `if` **guard** adds a condition — `case Point(x, y) if x == y:`. -Two traps that separate a careful answer from a naïve one: +Two points that separate a careful answer from a naïve one: - **A bare name is a capture, not a comparison.** `case foo:` does **not** test "is the value equal to `foo`?" — it matches _anything_ and rebinds `foo`, shadowing any outer variable. To match against an existing constant you need a **dotted name** (`case Color.RED:`) or a literal; this is why enums and module-qualified constants are the idiom. - It's most valuable for **decomposing complex, nested, heterogeneous data** — parsing an AST, dispatching on the shape of a JSON message, handling command objects. For a simple "one of N constants" branch, a plain `if/elif` or a dict dispatch table is clearer, and `match` is overkill. @@ -5074,7 +5100,7 @@ c = a print(a is c) # True - c is just another name for the same object ``` -**The one rule that matters in practice:** use `is` **only** for comparing against singletons — `None`, `True`, `False`, and sentinel objects. The idiom is `if x is None:`, never `if x == None:`, because a pathological `__eq__` could make `== None` lie, and `is None` is faster and unambiguous. For everything else — numbers, strings, containers — use `==`. +**The one rule that matters in practice:** use `is` **only** for comparing against singletons — `None`, `True`, `False`, and sentinel objects. The idiom is `if x is None:`, never `if x == None:`, because a badly written `__eq__` could make `== None` return the wrong answer, and `is None` is faster and unambiguous. For everything else — numbers, strings, containers — use `==`. **Where identity trips people up: caching makes `is` _appear_ to work on values, until it doesn't.** CPython pre-allocates and reuses certain immutable objects, so identity accidentally coincides with equality for them: @@ -5086,12 +5112,12 @@ print(x is y) # often True - compile-time string interning z = "".join(["h", "i"]); print(z is "hi") # often False - built at runtime, not interned ``` -These results are **implementation details** — they vary by Python version, by whether values are literals in the same code object, and between CPython/PyPy. Relying on `is` for value comparison produces bugs that pass in testing (small numbers, short literals) and fail in production (larger numbers, computed strings). That fragility is precisely _why_ the "`is` only for singletons" rule exists. +These results are **implementation details** — they vary by Python version, by whether the values are literals compiled together, and between CPython and PyPy. Relying on `is` for value comparison produces bugs that pass in testing (small numbers, short literals) and fail in production (larger numbers, computed strings). That fragility is precisely _why_ the "`is` only for singletons" rule exists. Senior-level footnotes: - **String interning** can be forced with `sys.intern(s)`, occasionally worth it when you compare many long strings repeatedly — interned strings compare by identity first, making `==` short-circuit. -- `id()` returns an object's identity (its address in CPython); it's unique only among _live_ objects, so a freed object's id can be reused. +- `id()` returns an object's identity (its memory address in CPython); it's unique only among _live_ objects, so a freed object's id can be reused. - A subtle gotcha: `float('nan') != float('nan')` is `True` (NaN is not equal to itself), yet `x = float('nan'); x is x` is `True` — a case where `is` and `==` genuinely diverge, and why containers use an identity check _before_ equality when searching. ## 129- What is monkey patching, and when is it appropriate? @@ -5102,7 +5128,7 @@ Senior-level footnotes: import some_library def patched(self, *args, **kwargs): - ... # your replacement behaviour + ... # your replacement behavior some_library.SomeClass.method = patched # swap the method out at runtime ``` @@ -5111,12 +5137,12 @@ some_library.SomeClass.method = patched # swap the method out at runtime - **Testing** — the single most defensible use. `unittest.mock.patch` is monkey patching with a safety net: it temporarily swaps a dependency for a mock/stub, then **restores the original automatically** when the test ends. Replacing a network call, a clock (`time.time`), or a database with a fake is standard practice. - **Hotfixing a third-party bug** you can't wait for upstream to fix or can't fork — patch the broken method in your own startup code as a stopgap. -- **Compatibility shims / backporting** — polyfilling a missing method so old and new versions of a library present the same interface. -- **Framework instrumentation** — some profilers, tracers, and greenlet libraries (e.g. `gevent`'s `monkey.patch_all()`) patch the standard library to inject their behaviour transparently. +- **Compatibility shims and backports** — adding a missing method so that old and new versions of a library present the same interface. +- **Framework instrumentation** — some profilers, tracers, and green-thread libraries (e.g., `gevent`'s `monkey.patch_all()`) patch the standard library to add their behavior transparently. **Why it's dangerous, and the senior's caution:** -- It's **action at a distance.** A patch applied in one module silently changes behaviour everywhere that code is used, so a reader of the affected class has no local indication that it was altered — nightmarish to debug. +- It causes **action at a distance.** A patch applied in one module silently changes behavior everywhere that code is used, so someone reading the affected class has no local indication that it was altered — which makes bugs very hard to track down. - It's **fragile against upgrades.** You're reaching into another library's internals; the next version can rename or restructure what you patched, breaking your code with no warning. - **Ordering and global state matter** — two patches of the same target, or a patch applied after the target is already used, produce order-dependent bugs. @@ -5124,7 +5150,7 @@ The rule a senior applies: **prefer the ordinary extension mechanisms first** ## 130- How do operators dispatch to dunder methods, and what is `NotImplemented`? -When you write `a + b`, Python doesn't have a single "add" — it dispatches to **`a.__add__(b)`**, and there's a fallback protocol that senior engineers are expected to understand, centred on the special sentinel **`NotImplemented`**. +When you write `a + b`, Python doesn't have a single built-in "add" — it dispatches to **`a.__add__(b)`**, and there's a fallback protocol that senior engineers are expected to understand, built around the special sentinel value **`NotImplemented`**. **The dispatch sequence for a binary operator `a + b`:** @@ -5139,25 +5165,28 @@ class Money: def __add__(self, other): if isinstance(other, Money): return Money(self.cents + other.cents) + if isinstance(other, int) and other == 0: + return self # lets sum() start from its default 0 return NotImplemented # let Python try other.__radd__ - __radd__ = __add__ # so sum([...]) and int+Money can work + __radd__ = __add__ # handles 0 + Money(...), which sum() does first def __repr__(self): return f"Money({self.cents})" print(Money(100) + Money(50)) # Money(150) +print(sum([Money(100), Money(50)])) # Money(150) ``` **The critical distinctions to get right:** - **`NotImplemented` is a singleton sentinel; `NotImplementedError` is an exception.** They are unrelated. Operator methods **return** `NotImplemented` to say "I don't handle this — try the other operand." Abstract methods **raise** `NotImplementedError` to say "a subclass must implement this." Confusing them is a classic mistake, and returning `NotImplementedError` (the class) from `__add__` silently breaks the fallback because it's a truthy object, not the sentinel. -- **Reflected methods** (`__radd__`, `__rmul__`, `__rsub__`, …) exist for every binary operator and are what make your type interoperate with built-in types on the left. `__rsub__` must remember the operands are swapped: `a - b` failing over calls `b.__rsub__(a)`, i.e. "compute a − b" from b's perspective. +- **Reflected methods** (`__radd__`, `__rmul__`, `__rsub__`, …) exist for every binary operator and are what make your type interoperate with built-in types on the left. `__rsub__` must remember that the operands are swapped: when `a - b` falls back, Python calls `b.__rsub__(a)`, which must compute `a − b` (not `b − a`). - **In-place operators** (`__iadd__` for `+=`, etc.) let mutable types mutate themselves and return `self`; if absent, `a += b` falls back to `a = a + b`, rebinding the name. - **Rich comparisons** (`__eq__`, `__lt__`, …) follow the same `NotImplemented` fallback, and returning `NotImplemented` from `__eq__` lets Python fall back to identity comparison rather than forcing a wrong answer. -The takeaway: implement operators so that they **return `NotImplemented` for types they don't recognise** rather than raising or guessing — that single discipline is what lets Python's reflected-operator machinery compose your types cleanly with each other and with the built-ins. +The takeaway: implement operators so that they **return `NotImplemented` for types they don't recognize** rather than raising or guessing — that one habit is what lets Python's reflected-operator machinery compose your types cleanly with each other and with the built-ins. ## 131- What is exception chaining, and what are `__context__`, `__cause__`, and `__suppress_context__`? -When a new exception is raised while another is being handled, Python **links the two** so the traceback tells the whole story. Formalised in PEP 3134, every exception carries three attributes that govern this: `__context__`, `__cause__`, and `__suppress_context__`. Understanding them is what lets you read — and control — multi-exception tracebacks. +When a new exception is raised while another is being handled, Python **links the two** so the traceback tells the whole story. Formalized in PEP 3134, every exception carries three attributes that govern this: `__context__`, `__cause__`, and `__suppress_context__`. Understanding them is what lets you read — and control — multi-exception tracebacks. **Implicit chaining.** Raising inside an `except` (or `finally`, or `with`) automatically sets the new exception's `__context__` to the one being handled: @@ -5192,7 +5221,7 @@ except ZeroDivisionError: Only the `RuntimeError`. `from None` sets `__cause__` to `None` and `__suppress_context__` to `True`, and the display rule then hides the context. -**The traceback display rule**, worth memorising: +**The traceback display rule**, worth memorizing: - If `__cause__` is present, **always** show it ("direct cause"). - Otherwise, show `__context__` **only if** `__suppress_context__` is `False` ("during handling"). @@ -5270,19 +5299,19 @@ warnings.simplefilter("error") # promote warnings to exceptions -> foo() now Two behaviors to know: -- By default each distinct warning is shown **once per location** and then suppressed (the `"default"` filter), which is why calling `foo()` a second time prints nothing new. +- By default, each distinct warning is shown **once per location** and then suppressed (the `"default"` action), which is why calling `foo()` a second time prints nothing new. Note also that `DeprecationWarning` is hidden by default unless it is triggered by code in `__main__` (PEP 565) or you run under a test runner such as pytest, which shows it. That is why library authors use `FutureWarning` for deprecations that end users must see. - The `"error"` filter turns warnings into real exceptions — and this is _precisely why_ the hierarchy roots at `Exception`: promoting a warning to an error is just raising it. Enabling `-W error` (or `filterwarnings = error` in pytest) is the standard way to make `DeprecationWarning`s **fail the build** in CI, catching deprecated usage before it breaks on an upgrade. When to use which: **raise an exception** for a problem the caller must handle right now; **emit a warning** for something that still works but shouldn't be relied on — deprecations, or suspicious-but-legal usage — leaving the final decision (ignore, show, or escalate to an error) to the user. ## 134- Walk through the full class-creation protocol: `__prepare__`, the metaclass, and how `__call__` controls instantiation -Question 120 covered _what_ a metaclass is; the senior follow-up is _what actually happens_ when a `class` statement runs, and how that differs from what happens when you later call the class to make an instance. These are two separate events driven by two different hooks, and conflating them is the classic mistake. +Question 120 covered _what_ a metaclass is; the senior follow-up is _what actually happens_ when a `class` statement runs, and how that differs from what happens when you later call the class to make an instance. These are two separate events driven by two different hooks, and mixing them up is the classic mistake. **Class-definition time — the four steps behind `class Foo(Base): ...`:** 1. **Determine the metaclass.** Python uses the explicit `metaclass=` if given, otherwise the most-derived metaclass among the bases, otherwise `type`. -2. **Prepare the namespace.** Python calls `metaclass.__prepare__(name, bases, **kwds)`, which returns the mapping used to execute the class body. The default is an ordinary `dict`, but returning a custom mapping lets you _record definition order_ or _forbid duplicate names_ — this is exactly how `EnumMeta` rejects two members with the same name. +2. **Prepare the namespace.** Python calls `metaclass.__prepare__(name, bases, **kwds)`, which returns the mapping used to execute the class body. The default is an ordinary `dict`, but returning a custom mapping lets you _record definition order_ or _forbid duplicate names_ — this is exactly how `EnumType` (formerly `EnumMeta`) rejects two members with the same name. 3. **Execute the class body** into that namespace — every method `def` and class variable becomes a key. 4. **Create the class object** by calling `metaclass(name, bases, namespace)`, which runs the metaclass's `__new__` (builds the class) then `__init__` (configures it). @@ -5290,14 +5319,14 @@ Question 120 covered _what_ a metaclass is; the senior follow-up is _what actual class OrderedMeta(type): @classmethod def __prepare__(mcs, name, bases, **kwds): - return {} # a real impl might return OrderedDict / a duplicate-guard + return {} # a real implementation might return a mapping that rejects duplicates def __new__(mcs, name, bases, ns): cls = super().__new__(mcs, name, bases, ns) cls._fields = [k for k in ns if not k.startswith("__")] return cls ``` -**Instance-creation time — a _different_ hook.** When you write `Foo(1, 2)`, Python does **not** call `Foo.__new__` directly. It calls `type(Foo).__call__` — i.e. the **metaclass's `__call__`** — and _that_ is what orchestrates the usual `instance = cls.__new__(cls, ...)` then `cls.__init__(instance, ...)` dance. Overriding the metaclass `__call__` is therefore the clean way to control instantiation itself — the correct way to build a true Singleton or an instance cache, avoiding the well-known pitfalls of hijacking `__new__`: +**Instance-creation time — a _different_ hook.** When you write `Foo(1, 2)`, Python does **not** call `Foo.__new__` directly. It calls `type(Foo).__call__` — i.e., the **metaclass's `__call__`** — and _that_ is what runs the usual sequence: `instance = cls.__new__(cls, ...)`, then `cls.__init__(instance, ...)`. Overriding the metaclass `__call__` is therefore the clean way to control instantiation itself — the correct way to build a true singleton or an instance cache, avoiding the pitfalls of overriding `__new__` (such as `__init__` running again on every call): ```python class Singleton(type): @@ -5313,26 +5342,26 @@ class Config(metaclass=Singleton): assert Config() is Config() # same object every time ``` -**The distinction to state crisply:** metaclass `__new__`/`__init__` run **once, when the class is defined**; metaclass `__call__` runs **every time you instantiate the class**. And as question 120 stressed, for almost all real needs the lighter hooks — `__init_subclass__` (react to subclassing) and `__set_name__` (descriptors learn their attribute name) — are the right tools; reach for `__prepare__` and a custom `__call__` only when you genuinely need to reshape the namespace or intercept construction. +**The distinction to state clearly:** metaclass `__new__`/`__init__` run **once, when the class is defined**; metaclass `__call__` runs **every time you instantiate the class**. And as question 120 stressed, for almost all real needs the lighter hooks — `__init_subclass__` (react to subclassing) and `__set_name__` (descriptors learn their attribute name) — are the right tools; reach for `__prepare__` and a custom `__call__` only when you genuinely need to reshape the namespace or intercept construction. ## 135- Reference cycles, `__del__`, `weakref`, and `gc.freeze()`: the practical garbage-collection questions -Question 11 laid out the two mechanisms — always-on reference counting plus a cyclic collector for the cycles refcounting can't see. The senior-level follow-ups probe the _interactions_ between those mechanisms and the tools you use to tame them. +Question 11 laid out the two mechanisms — always-on reference counting plus a cyclic collector for the cycles refcounting can't see. The senior-level follow-ups explore how those mechanisms _interact_ and the tools you use to control them. -**`__del__` and cycles — the historical trap.** A finaliser (`__del__`) that participates in a reference cycle used to be poison: before **PEP 442 (Python 3.4)** the collector couldn't decide a safe order to run finalisers in a cycle, so it gave up and dumped those objects into `gc.garbage`, leaking them forever. Since 3.4 finalisers _do_ run even inside cycles, but `gc.garbage` still exists and `__del__` remains unreliable for other reasons: its timing is tied to refcount reaching zero, it may **not run at all** at interpreter shutdown, and it can even _resurrect_ the object by creating a new reference to `self`. The rule: **never use `__del__` for resource cleanup** — use a context manager (`with`) or `weakref.finalize`, which is explicitly designed for this. +**`__del__` and cycles — the historical trap.** A finalizer (`__del__`) on an object in a reference cycle used to be a serious problem: before **PEP 442 (Python 3.4)** the collector couldn't decide a safe order to run finalizers in a cycle, so it gave up and dumped those objects into `gc.garbage`, leaking them forever. Since 3.4 finalizers _do_ run even inside cycles, but `gc.garbage` still exists and `__del__` remains unreliable for other reasons: its timing is tied to refcount reaching zero, it may **not run at all** at interpreter shutdown, and it can even _resurrect_ the object by creating a new reference to `self`. The rule: **never use `__del__` for resource cleanup** — use a context manager (`with`) or `weakref.finalize`, which is explicitly designed for this. -**`weakref` — references that don't keep objects alive.** A weak reference does _not_ increment the refcount, so it never prevents collection. Three canonical uses: +**`weakref` — references that don't keep objects alive.** A weak reference does _not_ increment the reference count, so it never prevents collection (see question 126). Three typical uses: - **Caches** that shouldn't pin their entries in memory — `weakref.WeakValueDictionary` / `WeakKeyDictionary`. -- **Breaking cycles** — e.g. a child holding a _weak_ back-reference to its parent so the pair can be collected by refcounting alone, no cyclic pass needed. +- **Breaking cycles** — e.g., a child holding a _weak_ back-reference to its parent, so the pair can be freed by reference counting alone, with no cyclic collection needed. - **Death callbacks** — `weakref.ref(obj, callback)` or `weakref.finalize(obj, cleanup)` to run code when the object goes away. **Tuning the collector.** It runs on **allocation-count thresholds** (`gc.get_threshold()`), not a clock. Practical levers: - `gc.disable()` in a **short-lived batch job** (the process exits before cycles matter) or in a **latency-sensitive request path** where you can't afford an unpredictable pause — then `gc.collect()` manually at a quiet moment. -- **`gc.freeze()` (Python 3.7+)** is the pre-fork server trick: call it _after_ loading your app but _before_ forking workers (gunicorn/uWSGI). It moves all currently-tracked objects into a permanent generation the collector ignores, so subsequent collections don't touch their GC headers — which keeps the shared, copy-on-write memory pages _clean_ across `fork()` instead of being dirtied by refcount/GC bookkeeping. This is the famous optimisation Instagram used to cut memory. +- **`gc.freeze()` (Python 3.7+)** is a trick for servers that fork worker processes (such as gunicorn or uWSGI): call it _after_ loading your app but _before_ forking the workers. It moves every object that exists at that point into a permanent generation that the collector ignores. Forked workers share the parent's memory pages until one of them writes to a page (copy-on-write); without `freeze()`, the collector's bookkeeping writes to those objects and forces each worker to make its own copy of the pages. Instagram popularized this technique to cut memory use. -**Diagnosing a leak in a long-running service:** reach for `tracemalloc` (snapshot + diff allocations by line), `gc.get_referrers(obj)` / `gc.get_objects()` to find what is holding an object alive, `objgraph` to visualise reference chains, and `gc.set_debug(gc.DEBUG_LEAK)` to log uncollectable objects. Remember the allocator subtlety from Q11: freeing objects returns memory to pymalloc's free lists, **not** always to the OS, so flat-but-high RSS is not necessarily a leak. +**Diagnosing a leak in a long-running service:** use `tracemalloc` (take snapshots and compare allocations by line), `gc.get_referrers(obj)` / `gc.get_objects()` to find what is holding an object alive, `objgraph` to visualize reference chains, and `gc.set_debug(gc.DEBUG_LEAK)` to log uncollectable objects. Remember the allocator subtlety from question 11: freeing objects returns memory to pymalloc's free lists, **not** always to the OS, so memory usage (RSS) that stays high but stable is not necessarily a leak. ## 136- What is `setup.py`, and how has Python packaging changed with `pyproject.toml`? @@ -5369,13 +5398,13 @@ dependencies = ["requests>=2"] mycli = "mypkg.cli:main" ``` -**Key points to land:** +**Key points to make:** -- **`setup.py` is not dead, but its role shrank.** It's now just one possible _configuration file_ for the setuptools backend, still useful for **compiled C/Rust extensions (`ext_modules`)** or genuinely dynamic metadata. But invoking it directly (`python setup.py install`, `sdist`, `bdist_wheel`) is **deprecated** — use `python -m build` and `pip` instead. `distutils` itself was **removed from the stdlib in Python 3.12**. +- **`setup.py` is not dead, but its role shrank.** It's now just one possible _configuration file_ for the setuptools backend, still useful for **compiled C/Rust extensions (`ext_modules`)** or genuinely dynamic metadata. But invoking it directly (`python setup.py install`, `sdist`, `bdist_wheel`) is **deprecated** — use `python -m build` and `pip` instead. `distutils` itself was **removed from the standard library in Python 3.12**. - **Build backends are pluggable:** `setuptools`, `hatchling`, `flit-core`, `pdm-backend`, `maturin` (Rust). Tools like Poetry, Hatch, PDM, and uv wrap this same PEP 517/518/621 flow. -- **Editable installs** (`pip install -e .`) — for live development — now work for `pyproject.toml`-only projects thanks to **PEP 660**, no `setup.py` required. +- **Editable installs** (`pip install -e .`), which let you edit the source and use the changes without reinstalling, now work for `pyproject.toml`-only projects thanks to **PEP 660** — no `setup.py` required. - **Entry points** declared here power both **console scripts** (CLI commands) and **plugin discovery** at runtime via `importlib.metadata.entry_points()`. -- Building produces an **sdist** (source) and a **wheel** (the installable built distribution — see the wheel-vs-egg question); publish with `twine upload dist/*`. +- Building produces an **sdist** (source distribution) and a **wheel** (the installable built distribution — see question 41); publish with `twine upload dist/*` (see question 150). The one-line takeaway: **declare metadata statically in `pyproject.toml` and let a PEP 517 backend build it; keep `setup.py` only for legacy projects or native extensions.** @@ -5386,7 +5415,7 @@ GraphQL is a **query language for APIs plus a runtime** that executes those quer **How it differs from REST:** - **Response shape is client-controlled.** REST returns a server-defined payload, which leads to **over-fetching** (you get fields you don't need) or **under-fetching** (you must call three endpoints to assemble one screen). A GraphQL query returns precisely the requested fields, and can traverse nested/related data in **one round trip**. -- **One typed schema, introspectable.** The schema (queries, mutations, subscriptions) is self-documenting and tooling-friendly (auto-complete, GraphiQL). +- **One typed schema that clients can inspect.** The schema (queries, mutations, subscriptions) is self-documenting and works well with tools (autocomplete, and the in-browser GraphiQL explorer). - **Trade-offs the interviewer wants to hear:** HTTP **caching is harder** (everything is a `POST` to one URL, versus REST's cacheable `GET`s and CDNs); **rate-limiting and observability** are trickier because one URL hides wildly different costs; and you must **defend against expensive queries** (depth/complexity limits, persisted queries) and the **N+1 problem**. **Serving it from Python.** The main libraries are **Strawberry** (modern, uses type hints/dataclasses, first-class FastAPI integration), **Graphene** (older, class-based), and **Ariadne** (schema-first SDL). A minimal Strawberry + FastAPI service: @@ -5420,40 +5449,40 @@ app.include_router(GraphQLRouter(schema), prefix="/graphql") Most application engineers do **inference, not training** — the job is to take a trained model and serve its predictions reliably inside a normal service. There are two integration patterns: -- **Call a hosted model API** (OpenAI, Vertex, Bedrock): no infrastructure, pay per call, but you inherit **network latency, rate limits, cost per request, and data-privacy** constraints (you're sending data to a third party). -- **Self-host the model**: load the weights **in-process** (PyTorch/`transformers`/scikit-learn) or behind a dedicated serving layer (**Triton, TorchServe, vLLM**, or a `FastAPI` wrapper). You control latency and data, but own the GPU/CPU capacity and ops. +- **Call a hosted model API** (such as OpenAI, Anthropic, Google Vertex AI, or Amazon Bedrock): no infrastructure and pay per call, but you inherit **network latency, rate limits, cost per request, and data-privacy** constraints (you're sending data to a third party). +- **Self-host the model**: load the weights **in-process** (PyTorch/`transformers`/scikit-learn) or behind a dedicated serving layer (**Triton, TorchServe, vLLM**, or a `FastAPI` wrapper). You control latency and data, but you own the GPU/CPU capacity and the operational work. -**The concern that trips people up in an async service:** model inference is **CPU/GPU-bound and synchronous**. In FastAPI/`asyncio`, running it directly in an `async def` handler **blocks the event loop** and stalls _every_ concurrent request. Offload it — `await loop.run_in_executor(pool, model.predict, x)` for a thread/process pool, or hand it to a **Celery worker** (exactly the `worker` mode this repo runs). This ties back to the GIL: heavy native libraries (PyTorch, NumPy) **release the GIL** during compute, so a thread pool genuinely helps; pure-Python pre/post-processing does not and needs processes. +**The concern that trips people up in an async service:** model inference is **CPU/GPU-bound and synchronous**. In FastAPI/`asyncio`, running it directly in an `async def` handler **blocks the event loop** and stalls _every_ concurrent request. Offload it — `await loop.run_in_executor(pool, model.predict, x)` for a thread or process pool, or hand it to a background task queue such as **Celery**. This ties back to the GIL: heavy native libraries (PyTorch, NumPy) **release the GIL** during compute, so a thread pool genuinely helps; pure-Python pre/post-processing does not and needs processes. **Lifecycle and operational concerns:** -- **Load the model once at startup**, keep it warm in memory (e.g. in the app lifespan), never per-request — loading weights is expensive. -- **Version the model artifact** independently of code; store it in a registry/S3, not the repo, and log which version served each prediction. -- **Throughput vs latency:** dynamic **batching** of requests, **caching** results/embeddings, streaming where possible. -- **Resilience for external calls:** timeouts, **retries with exponential backoff**, circuit breakers, and hard **cost/quota controls**. -- **Observability & correctness:** monitor latency and error rates, watch for **data/model drift**, log inputs/outputs with **PII care**, and pin the _preprocessing_ alongside the model so results stay reproducible. +- **Load the model once at startup** and keep it in memory (e.g., in the app's lifespan handler) — never load it per request, because loading weights is expensive. +- **Version the model artifact** independently of the code; store it in a model registry or object storage (such as S3), not in the repository, and log which version served each prediction. +- **Throughput vs. latency:** dynamic **batching** of requests, **caching** results and embeddings, and streaming responses where possible. +- **Resilience for external calls:** timeouts, **retries with exponential backoff**, circuit breakers, and hard **cost and quota controls**. +- **Observability and correctness:** monitor latency and error rates, watch for **data and model drift** (live data slowly diverging from the training data), log inputs and outputs while protecting personal data (**PII**), and version the _preprocessing_ code together with the model so results stay reproducible. The senior framing: treat the model as an **unreliable, expensive, versioned dependency** — isolate it behind a service boundary, keep it off the event loop, and wrap it in the same timeouts, retries, caching, and monitoring you'd give any external system. ## 139- What does a senior engineer need to know about building on LLMs (tokens, context windows, RAG, structured output, hallucination)? -At the API level an LLM is a **next-token predictor**: you send a prompt, it samples output tokens one at a time. The single most important mental model is that it is **stateless** — it has no memory between calls, so a "conversation" is an illusion you maintain by **resending the entire history** every request. +At the API level, an LLM is a **next-token predictor**: you send a prompt, it samples output tokens one at a time. The single most important mental model is that it is **stateless** — it has no memory between calls, so a "conversation" is an illusion you maintain by **resending the entire history** every request. -**Tokens and the context window.** Text is split into **tokens** (~4 characters / ¾ of a word in English). You are billed per **input + output** token and bounded by the **context window** — the maximum tokens for a single request (prompt _and_ completion). Count them with `tiktoken`; inputs that exceed the window must be **truncated, summarised, or chunked**. Runaway history is the usual cause of surprise bills and `context_length_exceeded` errors. +**Tokens and the context window.** Text is split into **tokens** (~4 characters / ¾ of a word in English). You are billed per **input + output** token and bounded by the **context window** — the maximum tokens for a single request (prompt _and_ completion). Count them with the provider's tokenizer or token-counting API (e.g., `tiktoken` for OpenAI models); inputs that exceed the window must be **truncated, summarized, or split into chunks**. Runaway history is the usual cause of surprise bills and `context_length_exceeded` errors. -**Sampling controls:** `temperature` / `top_p` govern randomness (near-0 for extraction/classification, higher for creative text), `max_tokens` caps output, `stop` sequences end generation, and `seed` gives _best-effort_ determinism. Streaming tokens back (SSE) is a UX necessity for anything long. +**Sampling controls:** `temperature` / `top_p` govern randomness (near-0 for extraction/classification, higher for creative text), `max_tokens` caps output, `stop` sequences end generation, and `seed` gives _best-effort_ determinism. Streaming tokens back as they are generated (usually via server-sent events, SSE) is essential for a good user experience with long responses. **Hallucination.** LLMs produce **fluent, confident, and sometimes false** output, and they cannot reliably tell when they're wrong. You mitigate — never fully eliminate — with grounding, asking for citations, constraining the output, external verification, and **human-in-the-loop** for high-stakes decisions. -**RAG (Retrieval-Augmented Generation) — the dominant grounding pattern:** **embed** your documents into vectors and store them in a **vector database** (pgvector, Pinecone, FAISS, Qdrant); at query time, embed the user's question, **retrieve the top-k most similar chunks**, and inject them into the prompt as context. This grounds answers in _your_ data, sidesteps the context-window and knowledge-staleness limits, and is how "chat over your documents" is built. +**RAG (Retrieval-Augmented Generation) — the dominant grounding pattern:** **embed** your documents into vectors and store them in a **vector database** (pgvector, Pinecone, FAISS, Qdrant); at query time, embed the user's question, **retrieve the top-k most similar chunks**, and inject them into the prompt as context. This grounds answers in _your_ data, works around both the context-window limit and the model's outdated training data, and is how "chat over your documents" is built. -**Structured output — don't parse prose.** For anything programmatic, force the model into a machine-readable shape via **JSON mode / function (tool) calling / a schema**, validate it (e.g. with **pydantic**), and retry on validation failure. Libraries like `instructor` and the OpenAI SDK's structured outputs do exactly this; orchestration frameworks (LangChain, LlamaIndex) help but can over-abstract — reach for them deliberately. +**Structured output — don't parse prose.** For anything programmatic, force the model into a machine-readable shape via **JSON mode, function (tool) calling, or a schema**, validate it (e.g., with **Pydantic**), and retry on validation failure. Libraries like `instructor` and the OpenAI SDK's structured outputs do exactly this; orchestration frameworks (LangChain, LlamaIndex) help but can over-abstract — reach for them deliberately. **Security and engineering discipline:** treat **model output and any retrieved/third-party content as untrusted input** — **prompt injection** is the LLM-era injection attack, so never let raw model output trigger privileged actions unchecked. Round it out with cost controls (cache, right-size the model, trim history), rate-limit/retry handling, and **evals** — regression tests for prompts, because a model or prompt change can silently degrade quality with no stack trace to warn you. ## 140- How do you actually test Python code — pytest fixtures, parametrization, and mocking? -Question 87 introduced `unittest`; in practice most modern teams reach for **`pytest`**, and a senior is expected to know _why_ and to test with discipline. Pytest replaces `unittest`'s ceremony (subclass `TestCase`, `self.assertEqual`) with **plain `assert`** — its rewritten assertion introspection prints the actual operands on failure — plus **fixtures**, **parametrization**, and a rich plugin ecosystem. +Question 87 introduced `unittest`; in practice, most modern teams use **`pytest`**, and a senior is expected to know _why_ and to test with discipline. Pytest replaces `unittest`'s boilerplate (subclassing `TestCase`, calling `self.assertEqual`) with **plain `assert`** statements — pytest rewrites them so that a failure shows the actual values being compared — plus **fixtures**, **parametrization**, and a rich plugin ecosystem. **Fixtures are dependency injection for tests.** A fixture is a function that builds something a test needs; the test requests it by naming it as a parameter. `yield` splits setup from teardown, and **scopes** (`function`, `class`, `module`, `session`) control how often it runs. Shared fixtures live in `conftest.py`, discovered automatically without imports: @@ -5464,7 +5493,7 @@ import pytest def db(): conn = connect() # setup yield conn # hand it to the test - conn.close() # teardown, even if the test fails + conn.close() # teardown, even if the test fails def test_user_count(db): # pytest injects the fixture by name assert db.count("users") == 0 @@ -5478,7 +5507,7 @@ def test_square(value, expected): assert square(value) == expected ``` -**Mocking — isolate the unit from the world.** Use `unittest.mock` (or the `mocker` fixture from `pytest-mock`) to replace slow/external dependencies. The single most common mistake is **patching where the object is _defined_ instead of where it is _used_** — you must patch the name in the module under test (`myapp.service.requests`, not `requests`). Configure behaviour with `return_value`/`side_effect` and assert interactions with `assert_called_once_with`: +**Mocking — isolate the unit from the world.** Use `unittest.mock` (or the `mocker` fixture from `pytest-mock`) to replace slow/external dependencies. The single most common mistake is **patching where the object is _defined_ instead of where it is _used_** — you must patch the name in the module under test (`myapp.service.requests`, not `requests`). Configure behavior with `return_value`/`side_effect`, and check interactions with `assert_called_once_with`: ```python def test_fetch(mocker): @@ -5488,7 +5517,7 @@ def test_fetch(mocker): m.assert_called_once_with("https://api/health", timeout=5) ``` -Know the vocabulary — **stub** (canned answers), **mock** (asserts on calls), **fake** (a working lightweight implementation, e.g. an in-memory DB) — and the `monkeypatch` fixture for env vars and attributes. Round it out with **markers** (`@pytest.mark.skip`, `xfail`, custom markers, `-k` selection), **coverage** (`pytest --cov`, aim >80%), and **property-based testing** (`Hypothesis`) to generate edge cases you'd never enumerate by hand. The senior framing: keep tests **fast, isolated, and deterministic**, follow the **test pyramid** (many unit, fewer integration, few E2E), and mock at the boundaries — not the internals. +Know the vocabulary — **stub** (returns canned answers), **mock** (lets you assert how it was called), **fake** (a working lightweight implementation, e.g., an in-memory database) — and the `monkeypatch` fixture for environment variables and attributes. Round it out with **markers** (`@pytest.mark.skip`, `xfail`, custom markers, `-k` selection), **coverage** (`pytest --cov`; a common target is 80% or more), and **property-based testing** (`Hypothesis`), which generates edge cases you'd never think to list by hand. The senior framing: keep tests **fast, isolated, and deterministic**, follow the **test pyramid** (many unit tests, fewer integration tests, and only a few end-to-end tests), and mock at the boundaries — not the internals. ## 141- What is Pydantic, and how does it differ from `dataclasses`? @@ -5512,9 +5541,9 @@ class User(BaseModel): User.model_validate({"id": "42", "name": "Ada", "email": "a@b.com"}) # id coerced "42"->42 ``` -**The crucial distinction from `dataclasses` (Q124):** a dataclass is _only_ a boilerplate reducer — it will happily store `User(id="not-an-int")` because annotations are **not enforced at runtime** (Q121). Pydantic exists precisely to **enforce** them: parse, validate, coerce, and serialize. Rule of thumb — **dataclasses for trusted internal data, Pydantic at the boundaries** where data arrives from outside. +**The crucial distinction from `dataclasses` (question 124):** a dataclass _only_ removes boilerplate — it will happily store `User(id="not-an-int")` because annotations are **not enforced at runtime** (question 121). Pydantic exists precisely to **enforce** them: parse, validate, coerce, and serialize. Rule of thumb — **dataclasses for trusted internal data, Pydantic at the boundaries** where data arrives from outside. -Points a senior should land: +Points a senior should make: - **v1 vs v2 matters.** Pydantic 2 rewrote the core in Rust (`pydantic-core`) for a large speedup and **renamed the API**: `.dict()`→`.model_dump()`, `.json()`→`.model_dump_json()`, `parse_obj`→`model_validate`, and the `@validator`/`@root_validator` decorators became `@field_validator`/`@model_validator`. Mixing v1 and v2 idioms is a common migration bug. - **`Field(...)`** adds constraints (`gt`, `max_length`, `default_factory`, aliases) and doc metadata that flows into the auto-generated **JSON Schema / OpenAPI**. @@ -5523,7 +5552,7 @@ Points a senior should land: ## 142- Beyond basic hints — what are `Protocol`, `TypeVar`/`Generic`, and how do you actually enforce types? -Questions 121 and 124 established that hints don't run and introduced `dataclasses`; the senior-level material is the **static** type system you build for tools like **mypy** and **pyright**, and the two features that make it powerful: **Protocols** and **generics**. +Question 121 established that type hints are not enforced at runtime. The senior-level material is the **static** type system you build for tools like **mypy** and **Pyright**, and the two features that make it powerful: **protocols** and **generics**. **`Protocol` — structural (duck) typing, statically.** Instead of requiring inheritance from a base class (nominal typing), a `Protocol` matches **any object that has the right shape**. This types Python's actual duck-typing idiom without forcing an inheritance hierarchy: @@ -5537,9 +5566,9 @@ def consume(src: Readable) -> bytes: # accepts files, sockets, BytesIO — any return src.read() ``` -Decorate with `@runtime_checkable` to allow `isinstance` against it (shallow — checks method presence only). +Decorate it with `@runtime_checkable` to allow `isinstance` checks against it (the check is shallow: it only verifies that the methods exist, not their signatures). -**`TypeVar`/`Generic` — parametric polymorphism.** A `TypeVar` lets a function or class be typed _in terms of_ the caller's type, so a container preserves element type instead of collapsing to `Any`. Python 3.12 (**PEP 695**) added clean built-in syntax: +**`TypeVar`/`Generic` — parametric polymorphism** (code that works with any type while keeping track of which type it is). A `TypeVar` lets a function or class be typed _in terms of_ the caller's type, so a container keeps its element type instead of collapsing to `Any`. Python 3.12 (**PEP 695**) added clean built-in syntax: ```python # classic @@ -5555,51 +5584,53 @@ class Stack[T]: def pop(self) -> T: ... ``` -TypeVars can be **bounded** (`TypeVar("T", bound=Number)`) or **constrained** (`TypeVar("S", str, bytes)`), and variance matters for correctness. Round out the toolbox: **`Optional[X]` / `X | None`**, **`Union` / `X | Y`**, **`Literal`** (exact values), **`Final`**, **`@overload`** (multiple typed signatures), **`ParamSpec`** (typing decorators that forward args), **`Self`**, **`Annotated`** (attach metadata — how FastAPI/Pydantic carry validation), and **`TypedDict`** for structured dicts. +TypeVars can be **bounded** (`TypeVar("T", bound=Number)`) or **constrained** (`TypeVar("S", str, bytes)`), and _variance_ (whether, say, a `Box[Dog]` may be used where a `Box[Animal]` is expected) matters for correctness. The rest of the toolbox: **`Optional[X]` / `X | None`**, **`Union` / `X | Y`**, **`Literal`** (exact values), **`Final`**, **`@overload`** (multiple typed signatures), **`ParamSpec`** (for typing decorators that forward arguments), **`Self`**, **`Annotated`** (attaches metadata — how FastAPI and Pydantic carry validation rules), and **`TypedDict`** for structured dicts. -**Enforcement is a _tooling_ decision, not a runtime one.** Run **mypy** or **pyright** in CI, ideally in **strict mode**, to catch type errors before runtime; guard import-cycle-only imports behind `if TYPE_CHECKING:`; and ship **stub files (`.pyi`)** for typing code you can't annotate inline. The payoff a senior emphasises: types are executable documentation that a machine verifies — most valuable on large, long-lived codebases and public APIs. +**Enforcement is a _tooling_ decision, not a runtime one.** Run **mypy** or **Pyright** in CI, ideally in **strict mode**, to catch type errors before runtime; put imports that are needed only for type hints (and would otherwise cause import cycles) behind `if TYPE_CHECKING:`; and ship **stub files (`.pyi`)** for code you can't annotate inline. The payoff a senior emphasizes: types are executable documentation that a machine verifies — most valuable on large, long-lived codebases and public APIs. ## 143- How do you find and fix a performance bottleneck in Python? -The senior answer begins with discipline, not tricks: **measure before you optimise.** Knuth's "premature optimization is the root of all evil" is the rule — profile to find the real hot spot, because it is almost never where you guess, and un-profiled "optimisations" trade readability for nothing. +The senior answer begins with discipline, not tricks: **measure before you optimize.** Knuth's "premature optimization is the root of all evil" is the rule — profile to find the real hot spot, because it is almost never where you guess, and "optimizations" made without profiling trade readability for nothing. **Know which tool answers which question:** - **`timeit`** — microbenchmark a single expression or snippet, correctly (many loops, best-of-N). - **`cProfile` + `pstats`** — deterministic, function-level profile: _which functions_ dominate cumulative/total time. The standard first pass. -- **`line_profiler`** — line-by-line timing _inside_ the hot function `cProfile` fingered. +- **`line_profiler`** — line-by-line timing _inside_ the hot function that `cProfile` identified. - **`py-spy`** — a **sampling** profiler that attaches to a **running production process without modifying or restarting it** — the tool for "prod is slow right now." Can emit flame graphs. - **Memory**: `tracemalloc` (stdlib, snapshot/diff allocations by line), `memory_profiler`, and `memray` / `scalene` (which profiles CPU _and_ memory together). -**Then optimise in order of leverage:** +**Then optimize in order of leverage:** -1. **Algorithm and data structure first** — the biggest wins are almost always Big-O. Swapping a repeated `x in some_list` (O(n)) for a `set` (O(1)) beats any micro-tuning (ties to Q90). -2. **Use the interpreter's fast paths** — built-ins and C-implemented libraries: `str.join` over `+=` in a loop, comprehensions over manual loops, and **vectorise with NumPy/pandas** to push loops into C (which also _releases the GIL_, Q1). -3. **Avoid repeated work** — hoist invariants out of loops, bind hot attribute/global lookups to locals, and **cache** (`functools.cache`/`lru_cache`, Q103). -4. **Reduce memory pressure** — `__slots__` (Q77) for many small objects, **generators** to stream instead of materialising large lists. -5. **Only then reach for heavy machinery** — `Cython`, `numba` (JIT), a C/Rust extension, or **concurrency** (choosing the right model per Q119, remembering the GIL means threads help I/O, processes help CPU). +1. **Algorithm and data structure first** — the biggest wins almost always come from better Big-O complexity. Replacing a repeated `x in some_list` (O(n)) with a `set` lookup (O(1)) beats any micro-tuning (see question 90). +2. **Use the interpreter's fast paths** — built-ins and C-implemented libraries: `str.join` instead of `+=` in a loop, comprehensions instead of manual loops, and **vectorize with NumPy/pandas** to move loops into C (which also _releases the GIL_; see question 1). +3. **Avoid repeated work** — move calculations that don't change out of loops, bind frequently used attributes and globals to local variables, and **cache** results (`functools.cache`/`lru_cache`; see question 103). +4. **Reduce memory pressure** — `__slots__` (question 77) for many small objects, and **generators** to stream data instead of building large lists. +5. **Only then reach for heavy machinery** — `Cython`, `numba` (a JIT compiler), a C/Rust extension, or **concurrency** (choosing the right model as in question 119, and remembering that, with the GIL, threads help I/O-bound work and processes help CPU-bound work). -The takeaway: **profile → fix the biggest thing → re-measure**, and stop when it's fast enough. A clear O(n) beating a clever O(n log n) with a huge constant is often the right engineering call. +The takeaway: **profile → fix the biggest problem → measure again**, and stop when it's fast enough. Remember that Big-O hides constant factors: for realistic input sizes, a simple O(n log n) solution can beat a clever O(n) one with a large constant, so measure rather than assume. -## 144- What are the security footguns every Python engineer must know? +## 144- What are the security pitfalls every Python engineer must know? -Security is where senior engineers earn their title, and Python has a specific set of traps worth naming precisely: +Security awareness is a core senior skill, and Python has a specific set of traps worth naming precisely: -- **Deserialising untrusted data = remote code execution.** **`pickle`**, `marshal`, and `yaml.load` can **execute arbitrary code** while loading. Never unpickle data you didn't produce; for untrusted input use **JSON**, and always **`yaml.safe_load`**. This is the single most dangerous, most-overlooked Python-specific footgun. -- **`eval` / `exec` / `compile` on any input derived from a user** is code injection. Almost always avoidable with `ast.literal_eval`, a real parser, or a lookup table. -- **Injection generally.** **SQL injection** — use **parameterised queries / bound parameters**, _never_ f-strings or `%` into SQL (an ORM helps but you can still foot-gun with raw SQL). **Command injection** — call `subprocess` with an **argument list and `shell=False`**, never `shell=True` on interpolated strings. **Path traversal** — validate/normalise paths against a base directory. -- **Randomness for security.** `random` is a **predictable PRNG** — never use it for tokens, passwords, or session IDs. Use the **`secrets`** module (`secrets.token_urlsafe`). Hash passwords with **bcrypt/argon2/scrypt**, never plain `md5`/`sha256`, and compare secrets with **`hmac.compare_digest`** to avoid timing attacks. -- **`assert` is stripped under `-O`.** Optimised bytecode (`python -O`) removes `assert` statements, so **never use `assert` for a security or validation check** in production code — the check silently vanishes. -- **Supply chain.** Pin dependencies, scan them (**`pip-audit`**, Safety, GitHub Dependabot), and beware **typosquatting** on PyPI. `pip install` can run arbitrary code from a malicious `setup.py` (Q136) — this is why static-metadata wheels are safer. -- **XML** — the stdlib parsers are vulnerable to entity-expansion/XXE attacks; use **`defusedxml`** for untrusted XML. +- **Deserializing untrusted data = remote code execution.** **`pickle`**, `marshal`, and `yaml.load` can **execute arbitrary code** while loading. Never unpickle data you didn't produce; for untrusted input use **JSON**, and always **`yaml.safe_load`**. This is the single most dangerous and most overlooked Python-specific pitfall. +- **`eval` / `exec` / `compile` on any input derived from a user** is code injection. It is almost always avoidable with `ast.literal_eval`, a real parser, or a lookup table. +- **Injection in general.** **SQL injection** — use **parameterized queries (bound parameters)**, _never_ f-strings or `%` formatting to build SQL (an ORM helps, but raw SQL can still be vulnerable). **Command injection** — call `subprocess` with an **argument list and `shell=False`**, never `shell=True` with interpolated strings. **Path traversal** — validate and normalize user-supplied paths against a base directory, so `../` can't escape it. +- **Randomness for security.** `random` is a **predictable pseudo-random number generator** — never use it for tokens, passwords, or session IDs. Use the **`secrets`** module (`secrets.token_urlsafe`). Hash passwords with **bcrypt/argon2/scrypt**, never plain `md5`/`sha256`, and compare secrets with **`hmac.compare_digest`** to avoid timing attacks. +- **`assert` is stripped under `-O`.** Optimized bytecode (`python -O`) removes `assert` statements, so **never use `assert` for a security or validation check** in production code — the check silently vanishes. +- **Supply chain.** Pin dependencies, scan them (**`pip-audit`**, Safety, GitHub Dependabot), and beware **typosquatting** on PyPI. `pip install` can run arbitrary code from a malicious `setup.py` when it builds from source (question 136) — this is why installing prebuilt wheels, which run no code at install time, is safer. +- **XML** — the standard library's XML parsers are vulnerable to several malicious-XML attacks (such as entity expansion, the "billion laughs" attack); use **`defusedxml`** for untrusted XML. -The mindset to convey: **treat every byte crossing a trust boundary as hostile** — deserialisation, subprocess arguments, SQL parameters, file paths, template inputs — and prefer the safe API by default. +The mindset to convey: **treat every byte crossing a trust boundary as hostile** — deserialization, subprocess arguments, SQL parameters, file paths, template inputs — and prefer the safe API by default. ## 145- How do you write your own context manager, and what's in `contextlib`? -Question 24 covered _using_ `with`; a senior is expected to _author_ context managers and know the toolkit. The protocol is two dunder methods: **`__enter__`** (runs on entry, its return value binds to `as x`) and **`__exit__`** (runs on exit — normally _or_ via exception — guaranteeing cleanup, which is why `with` beats manual try/finally and the unreliable `__del__` from Q135). +Question 24 covered _using_ `with`; a senior is expected to _author_ context managers and know the toolkit. The protocol is two dunder methods: **`__enter__`** (runs on entry, its return value binds to `as x`) and **`__exit__`** (runs on exit — normally _or_ via exception — guaranteeing cleanup, which is why `with` beats manual `try`/`finally` and the unreliable `__del__` from question 135). ```python +import time + class Timer: def __enter__(self): self.t0 = time.perf_counter() @@ -5609,9 +5640,9 @@ class Timer: return False # False => don't suppress an exception; True would swallow it ``` -The subtlety interviewers probe: **`__exit__` returning `True` suppresses the exception**; returning `False`/`None` lets it propagate. `__exit__` always runs, so it's the place to release locks, connections, or transactions. +The subtlety interviewers test: **`__exit__` returning `True` suppresses the exception**; returning `False`/`None` lets it propagate. `__exit__` always runs, so it's the place to release locks, connections, or transactions. -**The generator style is usually cleaner** — `@contextlib.contextmanager` turns a single-`yield` generator into a context manager, with setup before the `yield` and teardown after (wrap in try/finally so cleanup survives exceptions): +**The generator style is usually cleaner** — `@contextlib.contextmanager` turns a single-`yield` generator into a context manager, with setup before the `yield` and teardown after (wrap the `yield` in `try`/`finally`, or `try`/`except`, so cleanup survives exceptions): ```python from contextlib import contextmanager @@ -5627,33 +5658,33 @@ def transaction(conn): raise ``` -**`contextlib` toolbox worth naming:** `suppress(Exception)` (a clean "ignore this error"), `closing(obj)` (call `.close()` on exit), `redirect_stdout`/`redirect_stderr`, `nullcontext` (a no-op stand-in for optional CMs), and **`ExitStack`** — the power tool for entering a **dynamic/variable number** of context managers (e.g. opening a list of files) and unwinding them all correctly. And for `asyncio` (Q118) there's the async mirror: **`__aenter__`/`__aexit__`**, **`@asynccontextmanager`**, and **`async with`** — essential for async DB sessions and HTTP clients. +**`contextlib` toolbox worth naming:** `suppress(Exception)` (a clean "ignore this error"), `closing(obj)` (call `.close()` on exit), `redirect_stdout`/`redirect_stderr`, `nullcontext` (a do-nothing stand-in when a context manager is optional), and **`ExitStack`** — the power tool for entering a **variable number** of context managers (e.g., opening a list of files) and closing them all correctly. And for `asyncio` (question 118), there's the async mirror: **`__aenter__`/`__aexit__`**, **`@asynccontextmanager`**, and **`async with`** — essential for async DB sessions and HTTP clients. ## 146- What does a senior need to know about talking to a database (ORM vs Core, sessions, pooling, transactions, N+1)? -Most services live or die on their database layer, and **SQLAlchemy** is Python's dominant toolkit, so the interview centres there. The first distinction is **Core vs ORM**: **Core** is a Pythonic SQL expression language (you compose queries, rows come back as tuples/mappings); the **ORM** maps Python classes to tables and gives you objects with identity and change-tracking. They share one engine, and mixing them is normal. +Most services live or die by their database layer, and **SQLAlchemy** is Python's dominant toolkit, so interviews focus on it. The first distinction is **Core vs ORM**: **Core** is a Pythonic SQL expression language (you compose queries, rows come back as tuples/mappings); the **ORM** maps Python classes to tables and gives you objects with identity and change-tracking. They share one engine, and mixing them is normal. -**Connection pooling — the performance fundamental.** Opening a DB connection is expensive, so the **engine holds a pool** and hands out/returns connections. A senior can speak to `pool_size`, `max_overflow`, `pool_timeout`, and especially **`pool_pre_ping`** (validate a connection before use so a firewall-dropped or DB-restarted connection doesn't blow up the next request). Getting pool sizing wrong — too small starves throughput, too large exhausts the DB's connection limit — is a classic production incident. +**Connection pooling — the performance fundamental.** Opening a DB connection is expensive, so the **engine holds a pool** and hands out/returns connections. A senior can explain `pool_size`, `max_overflow`, `pool_timeout`, and especially **`pool_pre_ping`** (checks that a connection is still alive before using it, so a connection dropped by a firewall or a database restart doesn't make the next request fail). Getting pool sizing wrong — too small starves throughput, too large exhausts the DB's connection limit — is a classic production incident. -**The Session and the Unit of Work.** The ORM `Session` batches your changes and tracks objects via an **identity map** (one object per primary key per session). The distinctions that matter: +**The Session and the Unit of Work.** The ORM `Session` implements the _unit of work_ pattern: it collects your changes and writes them to the database together. It tracks objects through an **identity map** (one object per primary key per session). The distinctions that matter: - **`flush`** pushes pending SQL to the DB (so it's visible within the transaction) but doesn't end it; **`commit`** flushes _and_ commits the transaction; **`rollback`** discards it. -- Wrap work in a transaction and keep sessions **short-lived and request-scoped** — a long-lived shared session is a common bug (this repo's `RequestContextMiddleware` attaching `request.state.db` per request reflects that discipline). +- Wrap work in a transaction and keep sessions **short-lived and request-scoped** — a long-lived shared session is a common bug. The usual pattern in a web service is to open one session per request (for example, as a FastAPI dependency; see question 149) and close it when the response is sent. -**The N+1 query problem** (also raised for GraphQL, Q137): lazily loading a relationship inside a loop fires one query per parent → hundreds of round trips. Fix with **eager loading** (`selectinload`, `joinedload`). Other essentials: **Alembic** for schema **migrations**, **async** drivers (`asyncpg` + SQLAlchemy's async engine) so DB I/O doesn't block the event loop (Q118), and the maturity to know **an ORM isn't always right** — raw parameterised SQL (Q144) or a lighter query builder can be the better tool for complex analytical queries. +**The N+1 query problem** (also discussed for GraphQL in question 137): lazily loading a relationship inside a loop fires one extra query per parent row, which quickly adds up to hundreds of round trips. Fix it with **eager loading** (`selectinload`, `joinedload`). Other essentials: **Alembic** for schema **migrations**, **async** drivers (`asyncpg` plus SQLAlchemy's async engine) so database I/O doesn't block the event loop (question 118), and the judgment to know that **an ORM isn't always right** — raw parameterized SQL (question 144) or a lighter query builder can be the better tool for complex analytical queries. ## 147- How should logging be done in a production Python service? -The headline a senior states immediately: **use the `logging` module, never `print`.** `print` gives you no levels, no timestamps, no routing, and no way for operators to turn detail up or down. `logging` is a configurable framework built from four pieces: **Loggers** (what you call), **Handlers** (where records go — console, file, syslog, HTTP), **Formatters** (how they're rendered), and **Filters**. +The first thing a senior says: **use the `logging` module, never `print`.** `print` gives you no levels, no timestamps, no routing, and no way for operators to turn detail up or down. `logging` is a configurable framework built from four pieces: **Loggers** (what you call), **Handlers** (where records go — console, file, syslog, HTTP), **Formatters** (how they're rendered), and **Filters**. **The idioms that separate juniors from seniors:** -- **Get a module-level logger by name:** `logger = logging.getLogger(__name__)`. This builds the **dotted logger hierarchy** (`myapp.services.db`), so you can raise/lower verbosity per subsystem. Never log through the root logger directly. +- **Get a module-level logger by name:** `logger = logging.getLogger(__name__)`. This builds the **dotted logger hierarchy** (`myapp.services.db`), so you can raise or lower verbosity per subsystem. Never log through the root logger directly. - **Configure once, at the application entry point** — typically with `logging.config.dictConfig(...)`. **Libraries must not configure logging**; a well-behaved library only adds a `NullHandler` so it stays silent until the _application_ opts in. - **Use lazy `%`-style formatting**, not f-strings: `logger.info("user %s did %s", uid, action)`. The string is only interpolated **if that level is enabled**, saving work on suppressed `DEBUG` lines — and it keeps the message template stable for structured/aggregated logs. - **Log exceptions with the traceback:** inside an `except`, call `logger.exception("failed")` (or `logger.error(..., exc_info=True)`) to capture the stack. -**For real services, go structured.** Emit **JSON logs** (via `python-json-logger` or `structlog`) so a log aggregator (ELK, Loki, Datadog) can index fields, and attach a **correlation/request ID** — often propagated with **`contextvars`** (which, unlike thread-locals, work correctly across `async` tasks, Q118) — so you can trace one request across many log lines. Pitfalls to mention: **duplicate log lines** from adding handlers more than once or from **propagation** to ancestor loggers, and the cost of logging in tight loops. The takeaway: logging is an **operability feature** — design it so problems in production are diagnosable without a redeploy. +**For real services, go structured.** Emit **JSON logs** (via `python-json-logger` or `structlog`) so a log aggregator (ELK, Loki, Datadog) can index fields, and attach a **correlation (request) ID** — often propagated with **`contextvars`** (which, unlike thread-locals, work correctly across `async` tasks; see question 118) — so you can trace one request across many log lines. Pitfalls to mention: **duplicate log lines** from adding handlers more than once or from **propagation** to ancestor loggers, and the cost of logging in tight loops. The takeaway: logging is an **operability feature** — design it so problems in production are diagnosable without a redeploy. ## 148- Concurrency in practice: `concurrent.futures`, thread safety, and synchronization primitives @@ -5670,17 +5701,17 @@ with ThreadPoolExecutor(max_workers=8) as pool: **Thread safety — the misconception to correct:** the GIL does **not** make your code thread-safe. It guarantees a single _bytecode_ runs at a time, but a high-level operation like `counter += 1` is **read-modify-write across several bytecodes**, so two threads can interleave and lose an update. Any **check-then-act** or **compound** operation over shared mutable state is a race. -**The synchronization toolbox** (`threading`): **`Lock`** (mutual exclusion — always use `with lock:`), **`RLock`** (re-entrant, same thread can re-acquire), **`Semaphore`** (limit N concurrent), **`Event`** (signal between threads), **`Condition`** (wait-for-predicate), and **`Barrier`**. But the senior's preferred design is to **avoid shared mutable state altogether**: hand work between threads through a **`queue.Queue`**, which is itself thread-safe, turning locking into a producer/consumer pipeline. Know **deadlock** (two threads each holding a lock the other wants — prevent with consistent lock ordering) and **`threading.local`** for per-thread state. +**The synchronization toolbox** (`threading`): **`Lock`** (mutual exclusion: only one thread at a time; always use it as `with lock:`), **`RLock`** (re-entrant: the thread that holds it can acquire it again), **`Semaphore`** (allows up to N threads at once), **`Event`** (a signal between threads), **`Condition`** (wait until some condition becomes true), and **`Barrier`** (wait until N threads reach the same point). But the senior's preferred design is to **avoid shared mutable state altogether**: hand work between threads through a **`queue.Queue`**, which is itself thread-safe, turning locking into a producer/consumer pipeline. Know **deadlock** (two threads each holding a lock the other wants — prevent with consistent lock ordering) and **`threading.local`** for per-thread state. -**Two `ProcessPoolExecutor` gotchas** to mention: arguments and return values must be **picklable** (Q144's pickle caveats apply), and on Windows/macOS-spawn you must guard the entry point with **`if __name__ == "__main__":`** or you'll fork-bomb yourself. And the async equivalent (Q118) lives in `asyncio` — `asyncio.Lock`/`Semaphore`/`Queue` — which coordinate _tasks_ on one thread, not OS threads. +**Two `ProcessPoolExecutor` gotchas** to mention: arguments and return values must be **picklable** (see question 20), and you must guard the entry point with **`if __name__ == "__main__":`**. Without the guard, when workers are started with the "spawn" or "forkserver" method (the defaults on Windows and macOS, and on Linux since Python 3.14), each worker re-runs the main module's top-level code and tries to start workers of its own. And the async equivalent (question 118) lives in `asyncio` — `asyncio.Lock`/`Semaphore`/`Queue` — which coordinate _tasks_ on one thread, not OS threads. ## 149- What is ASGI vs WSGI, and how does a framework like FastAPI use dependency injection? -The foundational split is the **server-to-application interface**. **WSGI** is the classic **synchronous** standard (one request occupies a worker thread/process start to finish) behind Flask and traditional Django. **ASGI** is its **asynchronous** successor: it supports `async` handlers, long-lived connections (WebSockets, SSE), and lets **one worker juggle thousands of concurrent I/O-bound requests** on an event loop (Q118). **FastAPI** (on Starlette) is ASGI; you serve it with **`uvicorn`**, often managed by **`gunicorn`** running multiple uvicorn workers to use all CPU cores. This is exactly this repo's `api` mode. +The foundational split is the **server-to-application interface**. **WSGI** is the classic **synchronous** standard (one request occupies a worker thread/process start to finish) behind Flask and traditional Django. **ASGI** is its **asynchronous** successor: it supports `async` handlers, long-lived connections (WebSockets, SSE), and lets **one worker juggle thousands of concurrent I/O-bound requests** on an event loop (question 118). **FastAPI** (built on Starlette) is ASGI; you serve it with **`uvicorn`**, often managed by **`gunicorn`** running multiple uvicorn workers to use all CPU cores. -**Why FastAPI is the modern default:** it's **type-hint driven** — request/response bodies are declared as **Pydantic models** (Q141), so you get validation, serialization, and an **auto-generated OpenAPI/Swagger** spec for free from the same annotations. +**Why FastAPI is the modern default:** it's **driven by type hints** — request and response bodies are declared as **Pydantic models** (question 141), so you get validation, serialization, and an **auto-generated OpenAPI (Swagger)** spec for free from the same annotations. -**Dependency injection via `Depends`** is FastAPI's signature feature. You declare a dependency as a parameter; FastAPI **resolves and injects** it, caching per-request and unwinding any cleanup: +**Dependency injection via `Depends`** is FastAPI's signature feature. You declare a dependency as a parameter; FastAPI **resolves and injects** it, reuses the result within the same request, and runs any cleanup code after the response: ```python from fastapi import Depends, FastAPI @@ -5699,18 +5730,18 @@ async def read_user(uid: int, db=Depends(get_db)): return db.get(User, uid) ``` -The wins a senior highlights: dependencies are **reusable and composable** (auth, DB session, pagination, this repo's `auth.dependency`), and **overridable in tests** via `app.dependency_overrides` — which is why the codebase can swap in `BASF_FEDERATION_DEBUG_USER` to bypass auth under test. +The benefits a senior highlights: dependencies are **reusable and composable** (authentication, database sessions, pagination), and **overridable in tests** via `app.dependency_overrides` — for example, replacing the real authentication dependency with one that returns a fixed test user. -**The trap that ties it all together:** in an ASGI app, a **blocking call inside an `async def`** (a synchronous DB driver, `requests`, `time.sleep`, heavy CPU or model inference from Q138) **stalls the entire event loop and every concurrent request**. The fixes: use **async-native libraries** (`httpx`, `asyncpg`), or declare the handler as a plain **`def`** (FastAPI runs it in a threadpool), or offload heavy work to a **process pool or a Celery worker** (this repo's `worker` mode). Rounding out the picture: **middleware** (cross-cutting concerns like CORS and this repo's `RequestContextMiddleware`), **background tasks**, and the **lifespan** hook for startup/shutdown (opening the Mongo/Redis/S3 connections). This layered discipline — thin HTTP handlers, injected dependencies, business logic and I/O pushed into services — is precisely the router → service → helper architecture this project enforces. +**The trap that ties it all together:** in an ASGI app, a **blocking call inside an `async def`** (a synchronous DB driver, `requests`, `time.sleep`, or heavy CPU work such as the model inference from question 138) **stalls the entire event loop and every concurrent request**. The fixes: use **async-native libraries** (`httpx`, `asyncpg`), declare the handler as a plain **`def`** (FastAPI runs it in a thread pool), or offload heavy work to a **process pool or a Celery worker**. The rest of the picture: **middleware** (cross-cutting concerns such as CORS, request IDs, and timing), **background tasks**, and the **lifespan** hook for startup and shutdown (for example, opening database and cache connections). This layered discipline — thin HTTP handlers, injected dependencies, and business logic and I/O pushed into services — is the router → service → data-access architecture that keeps FastAPI services maintainable. ## 150- What are the concrete steps to publish a library to PyPI with pip (build/twine), Poetry, or uv? -Questions 92 and 136 covered the _concepts_ — distribution options and the `setup.py`→`pyproject.toml` shift. This is the _procedure_. The key insight that de-mystifies it: **all three tools do the same thing** — turn your project into the two standard artifacts (an **sdist** `.tar.gz` and a **wheel** `.whl`, Q41) and upload them to the same index, **PyPI**. They differ only in ergonomics. So learn the shared pipeline once, then the three command sets. +Questions 92 and 136 covered the _concepts_ — distribution options and the move from `setup.py` to `pyproject.toml`. This is the _procedure_. The key insight that makes it simple: **all three tools do the same thing** — turn your project into the two standard artifacts (an **sdist** `.tar.gz` and a **wheel** `.whl`; see question 41) and upload them to the same index, **PyPI**. They differ only in convenience and workflow. So learn the shared pipeline once, then the three sets of commands. **The shared pipeline (tool-agnostic):** -1. **Pick a unique, available name** — check `https://pypi.org/project//`; names are first-come and can't clash. -2. **Create accounts on both PyPI and TestPyPI**, enable 2FA, and create an **API token** (or set up Trusted Publishing, below). +1. **Pick a unique, available name** — check `https://pypi.org/project//`; names are first come, first served, and must be unique. +2. **Create accounts on both PyPI and TestPyPI**, enable two-factor authentication (2FA), and create an **API token** (or set up Trusted Publishing, described below). 3. **Lay out the project** with a `src/` layout and declare **static metadata in `pyproject.toml`** (PEP 621), plus a `README`, `LICENSE`, and a version: ```toml @@ -5740,17 +5771,17 @@ Questions 92 and 136 covered the _concepts_ — distribution options and the `se 5. **Verify** — check the metadata, and install the built wheel into a _fresh_ virtualenv to confirm it imports and runs. 6. **Upload to TestPyPI first, then to real PyPI.** -**A — The standards toolchain (`build` + `twine`, the "pip world"):** the most transparent, no extra framework. +**A — The standard toolchain (`build` + `twine`, the "pip world"):** the most transparent option, with no extra framework. ```bash python -m pip install --upgrade build twine -python -m build # -> dist/ (sdist + wheel) -python -m twine check dist/* # validate long-description/metadata -python -m twine upload --repository testpypi dist/* # dry-run on TestPyPI -python -m twine upload dist/* # publish to real PyPI +python -m build # -> dist/ (sdist + wheel) +python -m twine check dist/* # validate the long description and metadata +python -m twine upload --repository testpypi dist/* # trial run on TestPyPI +python -m twine upload dist/* # publish to the real PyPI ``` -Authenticate with an API token: username `__token__`, password `pypi-…` (set `TWINE_USERNAME`/`TWINE_PASSWORD` env vars or a `~/.pypirc`). +Authenticate with an API token: username `__token__`, password `pypi-…` (set the `TWINE_USERNAME`/`TWINE_PASSWORD` environment variables or use a `~/.pypirc` file). **B — Poetry (integrated deps + build + publish):** @@ -5762,41 +5793,43 @@ poetry config pypi-token.pypi pypi-xxxx poetry publish # (add --build to build+publish in one step) ``` -**C — uv (the fast, newer all-in-one, Rust-based):** +**C — uv (the newer, fast, all-in-one tool, written in Rust):** ```bash uv init --lib mylib # scaffold a library (pyproject + src layout) uv build # -> dist/ (sdist + wheel) -uv publish --publish-url https://test.pypi.org/legacy/ --token # test +uv publish --publish-url https://test.pypi.org/legacy/ --token # test uv publish --token # real PyPI (or set UV_PUBLISH_TOKEN) ``` **The cross-cutting essentials a senior stresses:** -- **A version is immutable and single-use.** Once `mylib 0.1.0` is uploaded, PyPI will **never** let you overwrite or re-upload that filename — you must **bump the version** for every release. Follow **SemVer**, keep **one source of truth** for the version (read it at runtime with `importlib.metadata.version("mylib")`), and remember you can _yank_ a bad release but not replace it. This is exactly why you **test on TestPyPI first** — so a mistake doesn't burn a real version number. -- **Trusted Publishing (OIDC) is the modern, secure CI path.** Instead of storing a long-lived API token in CI secrets, you configure PyPI to _trust_ your GitHub Actions/GitLab pipeline; it mints a **short-lived token per run**. The `pypa/gh-action-pypi-publish` action is the standard way. In real projects, **publishing runs in CI on a version tag** — build, test, then publish — not from a laptop. -- **The build backend is your choice** (Q136), independent of the publish tool: `setuptools`, `hatchling`, `flit-core`, `pdm-backend`, or `maturin` for Rust/native extensions. Poetry uses `poetry-core`; uv defaults to `hatchling` but honours whatever `[build-system]` you declare. -- **Metadata quality = a good PyPI page.** Fill `description`, `readme`, `license`, `classifiers`, `requires-python`, `[project.urls]`, and `[project.scripts]`/`[project.entry-points]` for CLIs and plugins. +- **A version is immutable and single-use.** Once `mylib 0.1.0` is uploaded, PyPI will **never** let you overwrite or re-upload that filename — you must **bump the version** for every release. Follow **SemVer**, keep **one source of truth** for the version (read it at runtime with `importlib.metadata.version("mylib")`), and remember you can _yank_ a bad release but not replace it. This is exactly why you **test on TestPyPI first** — so a mistake doesn't use up a real version number. +- **Trusted Publishing (OIDC) is the modern, secure CI path.** Instead of storing a long-lived API token in CI secrets, you configure PyPI to _trust_ your GitHub Actions or GitLab pipeline, and PyPI issues a **short-lived token for each run**. The `pypa/gh-action-pypi-publish` action is the standard way. In real projects, **publishing runs in CI on a version tag** — build, test, then publish — not from a laptop. +- **The build backend is your choice** (question 136), independent of the publishing tool: `setuptools`, `hatchling`, `flit-core`, `pdm-backend`, or `maturin` for Rust/native extensions. Poetry uses `poetry-core`; recent versions of `uv init` default to uv's own `uv_build` backend (earlier versions used `hatchling`), but uv respects whatever `[build-system]` you declare. +- **Good metadata makes a good PyPI page.** Fill `description`, `readme`, `license`, `classifiers`, `requires-python`, `[project.urls]`, and `[project.scripts]`/`[project.entry-points]` for CLIs and plugins. -The takeaway: the mechanics are identical — **`pyproject.toml` → build an sdist + wheel → upload to PyPI**. Pick **one** tool for ergonomics: **`build` + `twine`** for maximum transparency and control, **Poetry** for an integrated dependency-plus-publish workflow, or **uv** for speed and a single modern toolchain — then automate it in CI with **Trusted Publishing** and a tag-triggered release. +The takeaway: the mechanics are identical — **`pyproject.toml` → build an sdist + wheel → upload to PyPI**. Pick **one** tool based on the workflow you prefer: **`build` + `twine`** for maximum transparency and control, **Poetry** for an integrated dependency-plus-publish workflow, or **uv** for speed and a single modern toolchain — then automate it in CI with **Trusted Publishing** and a tag-triggered release. ## 151- Free-threaded Python (PEP 703, no-GIL) and subinterpreters (PEP 684 and 734): what actually changes for concurrency? -This is the deep end of the GIL thread that runs through Q1, Q2, Q23, and the "how do I pick a concurrency model" question (Q119). For roughly three decades the answer to "can Python threads run CPU-bound code in parallel?" was **no, because of the GIL**. Two _separate_ CPython efforts are now changing that answer, and the senior insight is that they attack the same problem **from opposite ends** — one removes the lock, the other replicates it. +This goes deeper into the GIL topic that runs through questions 1, 2, 23, and 119. For roughly three decades, the answer to "can Python threads run CPU-bound code in parallel?" was **no, because of the GIL**. Two _separate_ CPython efforts are now changing that answer, and the senior insight is that they attack the same problem **from opposite ends** — one removes the lock, and the other gives each interpreter its own. + +**Free-threaded CPython (PEP 703)** is a separate build of the interpreter — the "`t`" ABI (e.g., `python3.13t`) — that **removes the GIL entirely**, so multiple threads execute Python bytecode truly in parallel across cores. It was experimental in 3.13 and became officially supported, though still optional, in 3.14 (PEP 779). -**Free-threaded CPython (PEP 703)** ships as an experimental build starting in 3.13 — the "`t`" ABI (`python3.13t`) — and **removes the GIL entirely**, so multiple threads execute Python bytecode truly in parallel across cores. Keeping reference counting correct without one big lock required real machinery: **biased reference counting** (a cheap fast path for the object's owning thread plus an atomic shared counter), **immortal objects** (PEP 683 — `None`, `True`/`False`, small ints, and interned strings never touch their refcount), per-object locking, and a thread-safe variant of the pymalloc allocator (Q153). The costs are real: single-threaded code gets somewhat slower (the fast-path refcount checks and lost specialization from Q152), **C extensions must be recompiled and audited** for thread-safety because they can no longer assume the GIL serializes them, and it stays opt-in/experimental through 3.13–3.14. But it finally makes plain `threading` a legitimate answer for CPU-bound work. +Keeping reference counting correct without one big lock required real machinery: **biased reference counting** (a cheap fast path for the thread that owns the object, plus an atomic shared counter for other threads), **immortal objects** (PEP 683 — `None`, `True`/`False`, small ints, and interned strings never change their reference count), per-object locking, and a thread-safe allocator (see question 153). The costs are real: single-threaded code runs somewhat slower (roughly 5–10% in 3.14, and more in 3.13, where the specializing interpreter from question 152 was disabled on this build), and **C extensions must be recompiled and audited** for thread safety, because they can no longer assume the GIL runs them one thread at a time. But it finally makes plain `threading` a legitimate answer for CPU-bound work. -**Per-interpreter GIL and subinterpreters (PEP 684 for the runtime, PEP 734 for the stdlib API)** take the opposite route: **keep a GIL, but give each subinterpreter its own** so they stop contending. 3.12 moved most interpreter state to be per-interpreter; 3.13 exposes the `interpreters` module and a `concurrent.futures.InterpreterPoolExecutor`. Each subinterpreter is an isolated Python — its own imports, modules, and GIL — closer to `multiprocessing`'s isolation but living **in one process** (no `fork`, no child-process startup). You communicate over **queues/channels** that share only a narrow set of objects rather than passing arbitrary references around. +**Per-interpreter GIL and subinterpreters (PEP 684 for the runtime, PEP 734 for the stdlib API)** take the opposite route: **keep a GIL, but give each subinterpreter its own**, so they no longer compete for a single lock. Python 3.12 moved most interpreter state to be per-interpreter, and 3.14 exposes the feature in the standard library as the `concurrent.interpreters` module (PEP 734), plus `concurrent.futures.InterpreterPoolExecutor`. Each subinterpreter is an isolated Python — its own imports, modules, and GIL — closer to `multiprocessing`'s isolation but living **in one process** (no `fork`, no child-process startup). Subinterpreters communicate through **queues** that share only a narrow set of objects, rather than passing arbitrary references around. -So the mental model for choosing today refines Q119: **free-threading = shared memory, you manage the locks** (maximum performance, maximum footguns — data races are back), while **subinterpreters = isolated memory, you pass messages** (safer, but with marshaling overhead). For a service like this one, most parallelism is still I/O-bound and is pushed onto **Celery workers** (separate OS processes — the boring, robust answer), but the free-threaded build is why "just use a thread pool" may finally become viable for a CPU-bound endpoint without spawning processes. +So the mental model for choosing today refines question 119: **free-threading = shared memory, and you manage the locks** (maximum performance, but also maximum risk — data races are back), while **subinterpreters = isolated memory, and you pass messages** (safer, but data must be copied or serialized between interpreters). For a typical web service, most parallelism is still I/O-bound, and CPU-heavy work is pushed onto separate worker processes, such as **Celery workers** (the boring, robust answer). But the free-threaded build is why "just use a thread pool" may finally become viable for a CPU-bound endpoint without spawning processes. -**The trap** is assuming "no-GIL makes my existing threaded code faster/parallel." That is only true on the special build **and** only if every C-extension dependency supports it; on a normal build, nothing changes. Worse, removing the GIL **exposes latent races the GIL was accidentally hiding** — a non-atomic `counter += 1` on a shared object, or a check-then-act on a dict — so code that "worked" for years can start corrupting data. The synchronization discipline from Q148 (locks, `queue.Queue`, immutability) stops being optional. +**The trap** is assuming "no-GIL makes my existing threaded code faster/parallel." That is only true on the special build **and** only if every C-extension dependency supports it; on a normal build, nothing changes. Worse, removing the GIL **exposes latent races the GIL was accidentally hiding** — a non-atomic `counter += 1` on a shared object, or a check-then-act on a dict — so code that "worked" for years can start corrupting data. The synchronization discipline from question 148 (locks, `queue.Queue`, immutability) stops being optional. The takeaway: two PEPs, two philosophies — **PEP 703 removes the lock (shared-memory parallelism, bring your own synchronization)** and **PEP 684/734 replicates the lock per interpreter (message-passing isolation inside one process)** — and together they finally give CPython an in-process story for CPU-bound parallelism that used to force you into `multiprocessing`. ## 152- How does CPython execute bytecode? `dis`, the ceval loop, frames, and the specializing adaptive interpreter (PEP 659) -Q85 walked the pipeline (source → AST → bytecode → run) and Q86 covered `__pycache__`. This goes one level deeper: **what actually runs the bytecode**, and why it matters for performance and tracebacks. +Question 85 walked through the pipeline (source → AST → bytecode → run), and question 86 covered `__pycache__`. This goes one level deeper: **what actually runs the bytecode**, and why it matters for performance and tracebacks. **Compilation produces a code object.** The compiler turns each function into a **code object** reachable as `func.__code__`, holding `co_code` (the raw bytecode), `co_consts`, `co_varnames`, `co_names`, the required stack size, and flags. `dis.dis(func)` disassembles it into readable opcodes: @@ -5815,23 +5848,23 @@ dis.dis(f) **Frames tie calls together.** Each call creates a **frame** capturing the locals, the value stack, the instruction pointer (`f_lasti`), and a link to the caller (`f_back`). That `f_back` chain **is** your traceback, and it's what `sys._getframe()`, `sys.settrace`, debuggers, and profilers walk. Python 3.11 made frames dramatically cheaper — they live lazily as C structs on a per-thread data stack and are only materialized into full Python `frame` objects when something actually needs one — a major reason 3.11 was ~10–60% faster than 3.10. -**PEP 659 — the specializing adaptive interpreter (3.11+)** is the big modern idea: **quickening plus inline caching**. The interpreter observes which _types_ actually flow through a generic opcode and rewrites it in place into a **specialized** form — `BINARY_OP` on two ints becomes an int-add fast path; a `LOAD_ATTR` that keeps hitting the same class layout becomes a cached, guard-checked lookup. If an assumption breaks (a different type shows up), it **deoptimizes** back to the generic opcode. This is adaptive, JIT-_like_ behavior without a full JIT — and 3.13 added an experimental copy-and-patch JIT layered on top. +**PEP 659 — the specializing adaptive interpreter (3.11+)** is the big modern idea: **quickening** (rewriting instructions in place once code is "hot") plus **inline caches** (storing lookup results right next to the instruction). The interpreter observes which _types_ actually flow through a generic opcode and rewrites it in place into a **specialized** form — `BINARY_OP` on two ints becomes an int-add fast path; a `LOAD_ATTR` that keeps hitting the same class layout becomes a cached, guard-checked lookup. If an assumption breaks (a different type shows up), it **deoptimizes** back to the generic opcode. This is adaptive, JIT-_like_ behavior without a full JIT — and 3.13 added an experimental copy-and-patch JIT layered on top. -**Why a senior cares:** it explains _why_ micro-optimizations behave as they do — **monomorphic code** (the same types every call) specializes and runs faster than polymorphic code; it's why `dis` output on modern Python shows adaptive/specialized opcodes; and it demystifies tracebacks and tooling (they walk the frame chain). It also frames the free-threading trade-off in Q151, since specialization interacts with removing the GIL. +**Why a senior cares:** it explains _why_ micro-optimizations behave as they do — **monomorphic code** (the same types every call) specializes and runs faster than polymorphic code; it's why `dis.dis(func, adaptive=True)` on modern Python shows specialized opcodes; and it explains how tracebacks and tooling work (they walk the frame chain). It also explains part of the free-threading trade-off in question 151, since specialization interacts with removing the GIL. The takeaway: CPython compiles to a **code object**, then a **stack machine (`_PyEval_EvalFrameDefault`) executes it frame by frame**, and since 3.11 the **specializing adaptive interpreter** rewrites hot opcodes into type-specialized fast paths (deoptimizing when guards fail) — so `dis`, monomorphic code, and the frame chain are the three concepts to hold when reasoning about performance and stack traces. ## 153- A deeper look at CPython memory: `pymalloc` arenas, pools, and blocks, `tracemalloc`, and diagnosing leaks and fragmentation -Q11 covered the private heap plus refcount/GC, Q90/Q115/Q116 covered the `dict`/`list`/`tuple` layouts, and Q135 covered the cyclic collector. This is the **allocator layer underneath all of them** — the thing that decides where an object's bytes come from. +Question 11 covered the private heap, reference counting, and garbage collection; questions 90, 115, and 116 covered the `dict`/`list`/`tuple` layouts; and question 135 covered the cyclic collector. This is the **allocator layer underneath all of them** — the thing that decides where an object's bytes come from. **CPython does not call `malloc` per object.** It layers allocators: a raw domain (`PyMem_RawMalloc` → the system `malloc`), an object domain, and for small objects a dedicated allocator called **pymalloc (obmalloc)**. Objects **≤ 512 bytes** are served by pymalloc; anything larger goes straight to the system allocator. -**Arenas → pools → blocks.** pymalloc grabs memory from the OS in big **arenas** (256 KiB). Each arena is sliced into **pools** of 4 KiB (one OS page). A pool serves a **single size class** — block sizes are rounded up to a multiple of 8 (16 on 64-bit) bytes — so allocating an object is just a **free-list pop** of a fixed-size block: no syscall and minimal fragmentation _within_ a size class. +**Arenas → pools → blocks.** pymalloc requests memory from the OS in large **arenas** (1 MiB each on 64-bit builds since Python 3.10; 256 KiB before that and on 32-bit builds). Each arena is split into **pools** (16 KiB on 64-bit builds since 3.10; previously 4 KiB, one OS page). A pool serves a **single size class** — block sizes are rounded up to a multiple of 16 bytes on 64-bit builds (8 on 32-bit) — so allocating an object just takes a fixed-size block off a **free list**: no system call and minimal fragmentation _within_ a size class. (You can see the actual sizes on your build with `sys._debugmallocstats()`.) -**Why memory often doesn't go back to the OS.** An arena is only released when **all** of its pools and blocks are free. A single surviving object can pin an entire 256 KiB arena — the classic **fragmentation** surprise where RSS stays high long after you `del` most of your data. Freed blocks aren't returned to the OS either; they go on a free list to be reused by the same process. This is why "my long-running Python process never gives memory back" is usually _expected_ behavior, not a bug. +**Why memory often doesn't go back to the OS.** An arena is only released when **all** of its pools and blocks are free. A single surviving object can keep an entire arena alive — the classic **fragmentation** surprise where RSS stays high long after you `del` most of your data. Freed blocks aren't returned to the OS either; they go on a free list to be reused by the same process. This is why "my long-running Python process never gives memory back" is usually _expected_ behavior, not a bug. -**Immortal singletons tie in.** Small ints (−5..256), interned strings, and `None`/`True`/`False` are effectively permanent — and in 3.12+ literally _immortal_ (PEP 683), never refcounted or freed. That's by design: they're shared singletons (the identity gotchas of Q128), so pinning them costs nothing meaningful. +**Immortal objects fit in here too.** Small ints (−5 to 256), interned strings, and `None`/`True`/`False` are effectively permanent — and in 3.12+ literally _immortal_ (PEP 683): their reference counts never change, and they are never freed. That's by design: they're shared singletons (see the identity gotchas in question 128), so keeping them alive forever costs nothing meaningful. **Tooling to diagnose:** @@ -5850,21 +5883,21 @@ for stat in snap2.compare_to(snap1, "lineno")[:10]: print(stat) # the biggest growth by source line ``` -**"Leaks" in a garbage-collected language are almost always unintended references** (Q135): a module-level cache or list that only ever grows, an `lru_cache` with no `maxsize` pinning large results, a closure or callback that captured `self` and outlived it, or `__del__` on members of a reference cycle. The fixes are `weakref` (Q126), bounded caches, and explicitly breaking cycles. +**"Leaks" in a garbage-collected language are almost always unintended references** (question 135): a module-level cache or list that only ever grows, an `lru_cache` with no `maxsize` holding on to large results, a closure or callback that captured `self` and outlived it, or `__del__` on members of a reference cycle. The fixes are `weakref` (question 126), bounded caches, and explicitly breaking cycles. -**Repo-relevant:** a long-running FastAPI or Celery worker that slowly climbs in memory is the textbook case. Wrap a request or task in `tracemalloc` snapshots, look for unbounded caches, and remember that even after you fix the real leak, **RSS may not fall** because of arena fragmentation. +**A typical real-world case:** a long-running FastAPI or Celery worker whose memory use slowly climbs. Wrap a request or task in `tracemalloc` snapshots, look for unbounded caches, and remember that even after you fix the real leak, **RSS may not fall** because of arena fragmentation. The takeaway: small objects flow through **pymalloc's arena → pool → block hierarchy** (a free-list fast path, no per-object `malloc`), which is fast but means **memory returns to the OS only when a whole arena empties** — so high RSS is frequently fragmentation, while true leaks are lingering references best hunted with **`tracemalloc` snapshots plus `gc` referrer inspection** and fixed with `weakref` and bounded caches. ## 154- Generators as coroutines: `send`, `throw`, `close`, and `yield from` — and how they became `async`/`await` -Q36 and Q37 introduced generators and iterators; Q118 covered asyncio. This connects them, because the senior insight is that **native coroutines are generators grown up** — the `async`/`await` machinery is the generator protocol productized. +Questions 36 and 37 introduced generators and iterators; question 118 covered asyncio. This connects them, because the senior insight is that **native coroutines grew out of generators** — the `async`/`await` machinery is the generator protocol turned into dedicated syntax. **A generator is a two-way, resumable coroutine — not just a lazy iterator.** Beyond `next()`, `yield` is an _expression_ that can receive a value: - `gen.send(x)` — resume the generator, and the paused `yield` **evaluates to `x`** inside it. `send(None)` is equivalent to `next()`, and you must "prime" a generator with `next()`/`send(None)` before you can send a real value. - `gen.throw(exc)` — resume by **raising `exc` at the yield point**, which the generator can catch and handle. -- `gen.close()` — raise `GeneratorExit` at the yield point so `finally`/cleanup runs; this is exactly how `contextlib.contextmanager` (Q145) tears down after the `yield`. +- `gen.close()` — raise `GeneratorExit` at the yield point so `finally`/cleanup runs; this is exactly how `contextlib.contextmanager` (question 145) runs cleanup after the `yield`. ```python def averager(): @@ -5883,19 +5916,19 @@ print(a.send(20)) # 15.0 a.close() # raises GeneratorExit inside, runs cleanup ``` -**`yield from` (PEP 380)** delegates to a sub-generator: it transparently forwards `send`, `throw`, and `close` to the inner generator **and returns the inner generator's `return` value** (`result = yield from sub()`). Crucially it is _not_ just `for x in sub: yield x` — it wires up the full two-way protocol, which is precisely what you need to _compose_ coroutines out of smaller ones. +**`yield from` (PEP 380)** delegates to a sub-generator: it transparently forwards `send`, `throw`, and `close` to the inner generator **and returns the inner generator's `return` value** (`result = yield from sub()`). Crucially, it is _not_ just `for x in sub: yield x` — it wires up the full two-way protocol, which is precisely what you need to _compose_ coroutines out of smaller ones. -**The bridge to async/await.** Historically (3.4), asyncio coroutines literally _were_ generators — you wrote `@asyncio.coroutine` and drove them with `yield from`. Python 3.5 introduced **native coroutines** (`async def`/`await`) as a distinct type, but the underlying machinery is the same: `await` is the successor to `yield from` for driving awaitables, and the **event loop (Q118) repeatedly `.send()`s values into a coroutine and receives back the awaitable it suspended on**. The final result is carried out via `StopIteration.value`. In other words, the event loop is "just" a scheduler that drives many coroutine objects through `send`/`throw`. +**The bridge to async/await.** Historically (3.4), asyncio coroutines literally _were_ generators — you wrote `@asyncio.coroutine` and drove them with `yield from`. Python 3.5 introduced **native coroutines** (`async def`/`await`) as a distinct type, but the underlying machinery is the same: `await` is the successor to `yield from` for driving awaitables, and the **event loop (question 118) repeatedly `.send()`s values into a coroutine and receives back the awaitable it suspended on**. The final result is carried out via `StopIteration.value`. In other words, the event loop is "just" a scheduler that drives many coroutine objects through `send`/`throw`. Knowing they're generator-derived explains the rest of asyncio: **cancellation** is `coro.throw(CancelledError)` (the same mechanism as `gen.throw`), **cleanup on cancel** runs through `GeneratorExit`/`finally`, and a bare `yield` inside an `async def` produces an **async generator** consumed with `async for`. -**Gotchas:** forgetting to prime a `send`-based coroutine (you'll get `TypeError: can't send non-None value to a just-started generator`); swallowing `GeneratorExit` and then `yield`-ing again (Python raises `RuntimeError: generator ignored GeneratorExit`); and assuming `yield from`/`await` parallelize — they do not. One coroutine runs at a time on the loop; parallelism comes from `asyncio.gather`/`TaskGroup` (Q118, Q148). +**Gotchas:** forgetting to prime a `send`-based coroutine (you'll get `TypeError: can't send non-None value to a just-started generator`); swallowing `GeneratorExit` and then `yield`-ing again (Python raises `RuntimeError: generator ignored GeneratorExit`); and assuming `yield from`/`await` run things concurrently by themselves — they do not: awaiting a coroutine runs it to completion before the next line continues. Concurrency comes from scheduling several tasks, for example with `asyncio.gather` or a `TaskGroup` (questions 118 and 148), and even then only one coroutine runs at a time on the loop. -The takeaway: a generator is a **two-way, resumable coroutine** — `send` feeds values in, `throw` injects exceptions, `close` triggers cleanup via `GeneratorExit`, and `yield from` composes generators while forwarding all three plus the `return` value — and **`async`/`await` is that exact protocol productized**, with the event loop acting as the scheduler that `.send()`s coroutines forward. +The takeaway: a generator is a **two-way, resumable coroutine** — `send` feeds values in, `throw` injects exceptions, `close` triggers cleanup via `GeneratorExit`, and `yield from` composes generators while forwarding all three plus the `return` value — and **`async`/`await` is that same protocol turned into dedicated syntax**, with the event loop acting as the scheduler that `.send()`s coroutines forward. ## 155- Customizing classes without a metaclass: `__init_subclass__`, `__set_name__`, and the descriptor protocol in depth -Q120 and Q134 covered metaclasses, Q46 introduced descriptors, and Q32 covered `property`. The senior insight here is that **you rarely actually need a metaclass**: Python 3.6 (PEP 487) added two hooks that cover most real use cases with far less complexity and none of the metaclass-conflict pain. +Questions 120 and 134 covered metaclasses, question 46 introduced descriptors, and question 32 covered `property`. The senior insight here is that **you rarely need a metaclass**: Python 3.6 (PEP 487) added two hooks that cover most real use cases with far less complexity and none of the problems of conflicting metaclasses. **`__init_subclass__` — a classmethod on the _parent_ that runs once per subclass definition.** It lets a base class inspect, validate, or **register** its subclasses without a metaclass, and it receives keyword arguments passed in the class header. This is the clean way to build plugin registries or enforce that subclasses declare required attributes: @@ -5912,7 +5945,7 @@ class CsvLoader(PluginBase, key="csv"): # PluginBase.registry == {"csv": CsvLoader} ``` -**`__set_name__` — the hook that lets a descriptor learn its own attribute name.** When a class body assigns a descriptor to a name, Python calls `descriptor.__set_name__(owner, name)` **at class-creation time**, solving the old annoyance where a descriptor had no idea what attribute it was bound to (and you had to repeat the name). This is exactly how modern field libraries — SQLAlchemy 2.0 mapped columns, Pydantic-style models, dataclass-adjacent field systems — discover their attribute names: +**`__set_name__` — the hook that lets a descriptor learn its own attribute name.** When a class body assigns a descriptor to a name, Python calls `descriptor.__set_name__(owner, name)` **at class-creation time**, solving the old annoyance where a descriptor had no idea what attribute it was bound to (and you had to repeat the name). This is how many descriptor-based field libraries discover their attribute names: ```python class Field: # a data descriptor @@ -5924,14 +5957,14 @@ class Field: # a data descriptor setattr(obj, self.storage, value) ``` -**The descriptor protocol, deeper than Q46.** The three hooks are `__get__(self, obj, objtype)`, `__set__(self, obj, value)`, and `__delete__(self, obj)`. The critical distinction is **data descriptor** (defines `__set__` or `__delete__`) versus **non-data descriptor** (only `__get__`), because it drives attribute-lookup precedence: +**The descriptor protocol, in more depth than question 46.** The three hooks are `__get__(self, obj, objtype)`, `__set__(self, obj, value)`, and `__delete__(self, obj)`. The critical distinction is **data descriptor** (defines `__set__` or `__delete__`) versus **non-data descriptor** (only `__get__`), because it drives attribute-lookup precedence: -> **data descriptor on the type > instance `__dict__` > non-data descriptor > plain class attribute.** +> **data descriptor on the type > instance `__dict__` > non-data descriptor or plain class attribute.** That single rule explains a lot: **`property` is a data descriptor, so an instance attribute can't shadow it** (the property always intercepts). A plain **method is a non-data descriptor** — a function's `__get__` is precisely what **binds `self`** — which is why you _can_ override a method on a single instance by assigning to its `__dict__`. And **`functools.cached_property` is a non-data descriptor _on purpose_**: it computes once, writes the result into the instance `__dict__`, and thereafter the instance attribute shadows the descriptor so there's no recompute — a direct, practical consequence of the precedence rules. **Put together**, a typed/validating field implemented as a descriptor that learns its name via `__set_name__`, collected by a base class that uses `__init_subclass__`, is the _entire_ pattern behind modern data and ORM libraries — achieved with **no metaclass at all**. -**When you still genuinely need a metaclass (Q120/Q134):** when you must change what a class fundamentally _is_ — customizing `__prepare__` (the namespace used while the class body executes), altering the MRO, controlling `isinstance`/`__call__`, or transforming the class object before `__init_subclass__` even runs. Otherwise prefer the hooks: they **compose across multiple base classes** (metaclasses conflict and force you to write a combined metaclass), and they're far easier to read. +**When you still genuinely need a metaclass (questions 120 and 134):** when you must change what a class fundamentally _is_ — customizing `__prepare__` (the namespace used while the class body executes), altering the MRO, controlling `isinstance`/`__call__`, or transforming the class object before `__init_subclass__` even runs. Otherwise prefer the hooks: they **compose across multiple base classes** (metaclasses conflict and force you to write a combined metaclass), and they're far easier to read. The takeaway: reach for **`__init_subclass__`** (react to subclass creation — registries and validation) and **`__set_name__`** (let a descriptor discover its own attribute name) before ever writing a metaclass; combined with the **data vs non-data descriptor precedence** — the rule that explains why `property` beats an instance attribute while `cached_property` deliberately doesn't — these PEP 487 hooks cover the vast majority of "I thought I needed a metaclass" situations with dramatically less complexity.