Jump to Single Cycle CPU Documentation
Jump to Pipelined CPU Documentation
Jump to Cached CPU Documentation
- Yi Dong (yi.dong23@imperial.ac.uk)
- Seth Gobina (seth.gobina24@imperial.ac.uk)
- Zain Asif (zain.asif23@imperial.ac.uk)
- Mingze Chen (mingze.chen24@imperial.ac.uk)
(we also Implemented Cache, later)
The RISC-V CPU project was planned as a three-week incremental build, with the main focus on establishing a correct baseline design before introducing architectural complexity.
The first stage was the implementation of a single-cycle RV32I CPU, which served as the functional reference for the entire project. During this phase, the priority was correctness and clarity: defining clean module boundaries, validating control and datapath behavior, and ensuring that the processor could reliably execute the reference programs. This single-cycle design provided a stable foundation and a clear point of comparison for later extensions.
The second stage converted the single-cycle CPU into a five-stage pipelined architecture. Pipeline registers were introduced between IF, ID, EX, MEM, and WB, and the control logic was refactored accordingly. Data hazards were handled using forwarding and hazard detection, added incrementally and verified against previously working tests. The guiding principle was to reuse and adapt the existing single-cycle logic rather than redesigning the CPU from scratch.
After the pipelined CPU was stable, a simple cache layer was implemented by me independently in a short time. This extension focused on integration with the existing pipeline and correct stalling behavior, rather than aggressive optimization, and was built on top of the verified pipelined design.
As a team we decided to manage our repo in the following manner:
- Maintain a main branch containing the final and tested implementations of all CPU versions in their respective folders.
- The develop branch are used to store code and tb when they are verified
- we develop on seperate feature branches and combine them later for development
This approach provides a clear overview of our current progress, ensuring a clean and easily interpretable repository during examination.
| Module | Yi | Mingze | Seth | Zain |
|---|---|---|---|---|
| ALU.sv | L | |||
| ControlUnit.sv | C | L | ||
| DataMemory.sv | L | |||
| Extend.sv | L | |||
| InstrMem.sv | L | |||
| PCFlat.sv | L | |||
| RegFile.sv | L | |||
| top.sv | L |
Legend:
- L = Lead
- C = Contributor
- Performs arithmetic/logic operations
- 4 bit control signal to allow sufficient operations. Supports ADD, SUB, AND, OR, XOR, SLT, SLTU, SLL SRL, SRA operations. Allows for signed and unsigned. Used basic C++ testbench with functions to test all operations.
The Control Unit is the interpreter of the CPU. It fetches the 32-bit instruction and determines the control signals required by the rest of the architecture. These control signals dictate the outputs of multiplexers, the operations executed by the ALU, and which values are passed in as operands.
There are 3 flags that determine the instruction we will perform. They are extracted from the 32-bit instruction:
- Opcode – a broad identifier indicating the general instruction type.
- Funct3 and Funct7 – additional fields used to narrow down the exact instruction, since many share the same opcode.
- The remaining bits primarily encode immediate values.
To differentiate between all possible instructions, we used case() statements on the opcode, funct3, and funct7 fields.
Each control signal plays a specific role in directing data flow through the CPU:
- RegWrite – controls whether the Register File is written to.
- ALUControl – selects the ALU operation.
- ALUSrc – chooses the source of the ALU’s second operand.
- MemWrite – controls whether the Data Memory is written to.
- PCSrc – determines the next Program Counter value.
- ResultSrc – selects the data written back to the Register File.
- ImmSrc – identifies the immediate format so the ImmGen can reconstruct it.
- AddressingControl – determines how Data Memory constructs data for load and store instructions.
| Opcode | Type | RegWrite | ALUSrc | MemWrite | PCSrc | ResultSrc | ImmSrc | AddressingControl | Instructions |
|---|---|---|---|---|---|---|---|---|---|
| 0110011 | R-type | 1 | 0 | 0 | 00 | 00 | 000 | 000 | sub add sll slt sltu xor or and srl sra |
| 1100011 | Branch | 0 | 0 | 0 | 00/01* | 00 | 010 | 000 | beq bne |
| 0010011 | I-type ALU | 1 | 1 | 0 | 00 | 00 | 000 | 000 | addi slti sltiu slli xori ori andi srli srai |
| 1101111 | JAL | 1 | 0 | 0 | 01 | 10 | 011 | 000 | jal |
| 0000011 | Load | 1 | 1 | 0 | 00 | 01 | 000 | funct3 | lb lh lw lbu lhu |
| 0100011 | Store | 0 | 1 | 1 | 00 | 00 | 001 | funct3 | sb sh sw |
| 1100111 | JALR | 1 | 1 | 0 | 10 | 10 | 000 | 000 | jalr |
| 0110111 | LUI | 1 | 1 | 0 | 00 | 00 | 100 | 000 | lui |
Early versions of the CPU used word-addressable memory, which was sufficient for LW/SW but became cumbersome when supporting byte and halfword load/store instructions. To address this, I implemented a byte-addressable 128 KB data memory, allowing natural support for all RISC-V load and store variants. Memory accesses follow little-endian ordering, with byte, halfword, and word operations spanning 1, 2, or 4 consecutive bytes respectively, indexed using the lower 17 bits of the ALU address. Stores are synchronous, while loads are combinational and return ISA-correct values directly using a unified access_ctrl signal to handle size and sign extension. A $readmemh preload mechanism was added to support multiple test programs and reference workloads without modifying or recompiling the RTL.see more
For the program counter, we initially took the approach of separate components and a top-level interface with other components; however, in the end, this proved to be tedious and overly complicated. So, we turned to a flat implementation of the program counter. See implementation
- 32 bit registers with 2 read ports and one write port.
- Allows for asynchronous read and synchronous write on rising clock edge. x0 hardwired to 0. Simple unit testbench tested read and writing to registers, ensuring that they occur when expected (read whenever, write on rising clock edge).
The Immediate Generator receives:
- the full 32-bit instruction from ROM, and
- the ImmSrc signal from the Control Unit.
Depending on the immediate type, bits within the range [31:7] are arranged differently. The Generator extracts these fields and reconstructs the appropriate immediate value using a case() on the ImmSrc for each of the 5 different immediate types.
For example, for an I-Immediate, the Generator:
- Extracts bits [31:20].
- Sign-extends the 12-bit value to 32 bits.
This is correctly extracted according to the structure predefined for I-Immediates. Other immediate types follow similar patterns, each with its own bit arrangement and reconstruction rules.
| ImmSrc | Immediate Type | Bit Concatenation |
|---|---|---|
| 000 | I-type | 20 x Immediate[31] + Immediate[31:20] |
| 001 | S-type | 20 x Immediate[31] + Immediate[31:25] + Immediate[11:7] |
| 010 | B-type | 20 x Immediate[31] + Immediate[7] + Immediate[30:25] + Immediate[11:8] + 0 |
| 011 | J-type | 12 x Immediate[31] + Immediate[19:12] + Immediate[20] + Immediate[30:21] + 0 |
| 100 | U-type | Immediate[31:12] + 12 x 0 |
- Implemented single cycle datapath by connecting together the individual components.
- Instantiated and wired together the PC, Instruction Memory, Register File, Immediate Extend unit, ALU, Data Memory, and writeback modules. Implemented ALUSrc and ResultSrc multiplexers, PC source selection for normal, branch, JAL and JALR execution, and exposed internal signals (pc, instr, alu_result, a0) for easier debugging and verification.
add sub sll slt sltu xor srl sra or and
beq bne
addi slli slti sltiu xori srli srai ori andi lb lh lw lbu lhu jalr
jal
sb sh sw
lui
bltbgebgeubltuhave only been implemented in the pipelined version. Single cycle only implementsbeqandbne.
| Module | Yi | Mingze | Seth | Zain |
|---|---|---|---|---|
| fetch.sv | L | |||
| Decode.sv | L | L | ||
| EXECUTE_STAGE.sv | L | |||
| MEM_STAGE.sv | L | |||
| WB_STAGE.sv | L | |||
| IF_ID.sv | L | |||
| ID_EX.sv | L | |||
| EX_MEM.sv | L | |||
| MEM_WB.sv | L | |||
| HazardUnit.sv | L | |||
| PipelineTop.sv | L |
Legend:
- L = Lead
- C = Contributor
mux.sv is a 4 input multiplexer used as 3 instances in the architecture.
Legend: L = Lead C = Contributor
The pipeline of each stage is the one to its left.
We decided on this convention since we thought it would be easier to reason between stalling and flushing. Stalling implies that the inputs to the stage should remain unchanged, whereas flushing entails zeroing the inputs to the stage. As such, it made sense to define each pipeline of a stage to to the left of its corresponding stage.
Additions were made to the alu block. It implements the rest of the branch type instructions, by setting the Zero flag high whenever a branch test is passed for example, for blt, if SrcA < SrcB, Zero is high. This means that if the current instruction in execute stage happens to be a branch type instruction and the test is passed, BranchE will be high from the control unit, and so will Zero from the ALU. As such, ALUSrc will be high, and a branch will be effected.
The improvements in the control unit module improve efficiency, readability, and robustness of control logic. This includes better handling of instruction types, streamlined input processing, and the introduction of new control signals for enhanced functionality. key differences:
-
Rename signals like
RegWrite,ALUControl,MemWrite, etc. toRegWriteD,ALUControlD,MemWriteD, to prepare for pipeline integration -
Remove the
zeroinput and reducefunct7to a single logic bit, streamlining inputs -
Introduce new output signals like
JumpD,BranchD, andJALRInstrD. These signals add more control functionality, offering finer control over jump, branch, and JALR (Jump and Link Register) instructions. -
Enhances the control logic for B-type instructions by including additional operations (
blt,bge,bltu,bgeu). case structure simplified for better readability and maintenance. -
default case added in the main
casestatement, ensuring that all control signals are explicitly set to a known state when an unrecognized opcode is encountered. -
more cleanly structured with consistent indentation and improved commenting, improving code readability and maintainability.
The hazard unit produces StallFetch, StallDecode, FlushExecute, FlushDecode. These are inputs to the relevant pipelines for those stages that need to be flushed or stalled. Inside the pipelined, when stall signal is high, the signals at the pipeline's inputs are not passed to the outputs, while when flush signal is high, the outputs are low.
Each pipeline is in its own module, and those that are flushed / stalled at some point have internal signals to control that. Each stage is in its own module; the inputs to the module are those that are actually used for computing some value in that stage, while those that aren't used are connected directly to the next pipeline in the top level module.
The decode stage is in essence a top level module for the control unit, register file, and immediate generator, and brings together all of these into a singular module. It carries out the decoding of instructions, produces control signals, performs the extraction of data from the Register File, and formulates an immediate through the ImmGen module. All of these data are then forwarded to the ID-EX register, where the signals from the Hazard module determine how they are passed along. The decode module also passes signals from the IF-ID register, but these signals are not directly involved in the workings of the module.
- Pipeline registers were implemented between each stage to seperate datapath and control signals across cycles.
- Created IF/ID (Instruction Fetch/Decode), ID/EX (Decode/Execute), EX/MEM (Execute/Memory), and MEM/WB (Memory/Writeback) register modules to latch both data signals and control signals. Each register supports stall (holding state) and flush (inserting bubble) so hazard unit can handle data and control hazards.
The Hazard unit allows for the pipelined CPU to be able to perform instructions correctly without incurring delays for some special cases to ensure that it is as efficient as possible.
There are 3 different cases that we encountered that poses a challenge to pipelining and may result in an error if not taken care of which are the following:
- When we use a register as an operand that was written to in the previous cycle. (RAW hazard)
- When we have a branch instruction where we only know if we jump or not two cycles later in the execute stage.
- Load instructions where it takes an extra cycle to load data.
Thus, to solve these possible issues that the processor might encounter with pipelining we implement the following in our design:
- Forwarding: allows the value of a register to be used in an operation right after it was written without having to wait for it to go through all the pipelining stages.
- Stalling: Stalling a stage means to maintain its state, so the inputs to the stage should not change on the next clock cycle. This allows for load instruction to have its values from memory to be loaded into the writeback stage so that it can be forwarded onto the execute stage.
- Flushing: This resets the output of the pipeline flip-flops; This is very useful, for example in the case of branch, we do not know whether to jump or not until the branch instruction is in the execution stage. That means the next instruction in the instruction memory would be loaded onto the decode stage. This would create an error if the jump actually occurs, therefore we need to flush the decode stage when jump happens as if the instruction had never been loaded to the decode stage.
All three solutions/operations mentioned above are implemented in our pipelined CPU. Each operation may be used individually or simultaneously for specific cases/instruction. The control signal for forwarding, flush and stall are all produced/controlled by the hazard unit. RAW hazards are mitigated by forwarding from the Writeback or memory stages into the execute stage. If the current instruction in execute stage has a source register that's the same as the destination register as an instruction currently in writeback or memory stage, we forward data from writeback or memory stages respectively. We also only forward data from instructions that were going to write to a register. The zero register is never forwarded because it never has meaningful data being written to it (as it is hardwired to zero).
Lw issue is solved by stalling the decode and fetch stages. As such, we must flush the execute stage to prevent incorrect data from propagating forward.
If a control hazard is detected, the execute and decode stages are flushed (2 instructions after branch instruction are flushed) before moving to correct instruction.
the cache is solely developed by yi due to lacking time, see his repo
- There is a script located in the root of the directory called start.sh
- This links to a master script contained within /testing/Master_test
- Executing the script displays a menu that enables the execution of various programs on the specified CPU.
In order to view values in a particular register of the CPU, we added a signal testRegAddress which is controlled at the top level module, and outputs data from a given register at the signal testRegData. This allows use to use register data to view outputs on vbuddy, which is useful for pdf plots and f1 program.
- Move into the
testing/Master_testdirectory - Choose the
cpu_tb.cpptest bench using single cycle, andpipe_cpu_tb.cppif testing pipelined cpu - Both test benches write to CPU.vcd
Below is a code snippet of the test bench.
int main(int argc, char **argv, char **env)
{
int simcyc;
int tick;
char prog = argv[argc - 1][0];
Verilated::commandArgs(argc, argv);
// init top verilog instance
Vtop *top = new Vtop;
// init trace dump
Verilated::traceEverOn(true);
VerilatedVcdC *tfp = new VerilatedVcdC;
top->trace(tfp, 99);
tfp->open("CPU.vcd");
// init Vbuddy
if (vbdOpen() != 1)
{
return (-1);
}
vbdHeader("CPU");
vbdSetMode(1);
// intialise
top->clk = 1;
top->rst = 1;
step_cycle(top);
top->rst = 0;
// run simulation for MAX_SIM_CYC clock cycles
for (simcyc = 0; simcyc < MAX_SIM_CYC; simcyc++)
{
// dump variables into VCD file and toggle clock
for (tick = 0; tick < 2; tick++)
{
tfp->dump(2 * simcyc + tick);
top->clk = !top->clk;
top->eval();
}
// Test data
if (simcyc > 37530) // gaussian=123705, noisy=204890, triangle=316018, sine=37530
{
vbdPlot(top->a0, 0, 255);
vbdBar(top->a0 & 0xFF);
vbdCycle(simcyc);
}
// either simulation finished, or 'q' is pressed
if (Verilated::gotFinish() || vbdGetkey() == 'q')
exit(0);
}
vbdClose();
tfp->close();
exit(0);
}- The master script written in /testing/Master_test automatically configures the outputs depending on the program being run by passing arguments on execution
The tests for both single cycle and the pipelined CPU were written up here and here respectively using programs specified in the testing folder.
The following videos demonstrate the F1 program's functionality on a pipelined CPU with both data memory cache and instruction memory cache.the best way to test is to go to our clean version at Pureversion

