Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

AlexNet Final Project

Overview

This project is a hardware implementation of a reduced AlexNet-style convolutional neural network written in SystemVerilog. The goal was to build an end-to-end CNN inference pipeline that stays faithful to the overall AlexNet structure while remaining practical to simulate, debug, and reason about in a class project setting.

The design includes:

  • five convolutional layers
  • three max-pooling layers
  • one flatten stage
  • three fully connected layers
  • a top-level finite state machine to sequence inference

This is intentionally a "lite" AlexNet rather than a full reproduction of the original network. The layer ordering and high-level structure match the classic architecture, but the channel counts and output sizing were reduced to make simulation feasible. Note this project was the final project for CompSci-151 Digital Logic Design at UCI.

Background

AlexNet is a convolutional neural network architecture developed for image classification. The original model was designed for large-scale classification tasks and is much larger than what is implemented here. This project focuses on the hardware structure and dataflow of an AlexNet-style model rather than recreating the full production-scale network.

Instead of targeting one thousand ImageNet classes with a large trained parameter set, this implementation uses a smaller internal configuration and produces 10 output logits.

High-Level Architecture

The top-level inference flow is:

  1. conv1 -> pool1
  2. conv2 -> pool2
  3. conv3 -> conv4 -> conv5 -> pool3
  4. flatten
  5. fc1 -> fc2 -> fc3

At the end of inference, the design reports:

  • logits: the final class scores
  • pred_class: the index of the maximum logit
  • cycle_count: the number of cycles used for the inference

The top-level controller is implemented in SourceFiles/alexnet_top.sv, where a finite state machine launches each stage one at a time and waits for its done signal before moving on.

Module Descriptions

alexnet_top.sv

This is the top-level integration module. It stores the input image and network weights in internal arrays, instantiates every layer in the network, flattens the final pooled tensor, tracks total cycle count, and computes the final predicted class using an argmax over the output logits.

It also contains the main control FSM that sequences the entire inference pass from the first convolution to the final fully connected layer.

conv_layer.sv

This module implements a convolution layer using an internal finite state machine. It iterates over output channels, output coordinates, input channels, and kernel positions in a mostly sequential manner.

The implementation handles padding explicitly and accumulates the weighted sum for one output element at a time. To improve performance, it uses mac4_unit to evaluate four multiply-accumulate terms per cycle instead of only one.

maxpool_layer.sv

This module implements max pooling sequentially. For each output location and each channel, it scans the full pooling window, tracks the largest value seen so far, and writes the final maximum to the output feature map.

This design favors clarity and correctness over aggressive parallelism.

fc_layer.sv

This module implements a fully connected layer with a similar control style to conv_layer.sv. The layer computes one output neuron at a time and accumulates a weighted sum across the input vector.

Like the convolution layer, it uses mac4_unit so that four multiply-accumulate terms can be processed in parallel during each accumulation step.

mac4_unit.sv

This is a combinational helper block that performs four activation-weight multiplications in parallel, sums the products, and adds the result to the incoming accumulator value.

This module is the main datapath optimization used to reduce inference latency.

quantize_relu.sv

This helper module applies a right shift for simple quantization, then applies ReLU, and finally clamps the result into the 8-bit activation range used by the rest of the design.

cnn_pkg.sv

This package acts as the centralized configuration file for the design. It defines:

  • shared numeric types
  • utility functions
  • image dimensions
  • per-layer dimensions
  • flattened vector length

Keeping all sizing information in one package makes the rest of the modules easier to parameterize and maintain.

Data Representation

The project uses:

  • 8-bit unsigned activations
  • 8-bit signed weights
  • 32-bit signed accumulators

The configured input image size is:

  • 227 x 227 x 3

The fully connected output layer produces:

  • 10 output logits

This makes the implementation much smaller than full AlexNet, which was one of the major design choices for keeping the project simulation-friendly.

Design Decisions

Several major design choices shaped the final implementation.

1. Reduced channel sizes

The original AlexNet architecture is too large to simulate comfortably in a course-project workflow. To keep the design debuggable and simulation-friendly, the spatial structure and layer ordering were preserved while the channel counts were reduced.

This keeps the model recognizable as AlexNet-style hardware without requiring the scale of the original network.

2. FSM-driven layer control

Each major layer is driven by an internal finite state machine. This makes the logic easier to debug because the computation unfolds in a clear sequence of states instead of a deeply parallel datapath with more complicated control.

The tradeoff is speed. A purely sequential or mostly sequential architecture is much easier to reason about, but it can be slow.

3. Hybrid sequential / parallel execution

To improve performance without making the design dramatically harder to debug, the implementation uses a hybrid approach. Most of the control is still sequential, but the datapath does four multiply-accumulate operations at a time by using mac4_unit.

This provided a large performance improvement while keeping the architecture understandable.

4. Internal array storage

The input image, weights, and feature maps are stored in internal arrays rather than behind a separate memory interface. This saved development time and simplified testing.

The tradeoff is that the design is less realistic as a deployable accelerator and more tightly tied to simulation.

5. No memory reuse

Feature maps are stored in dedicated arrays between layers rather than reusing a smaller shared memory buffer. This simplifies the interfaces between modules and avoids extra address/control complexity.

The tradeoff is higher memory usage.

6. Cycle count tracking

The top-level module tracks cycle count so that design changes can be measured directly. This made it possible to compare the original sequential version against the optimized version with mac4_unit.

Performance

One of the biggest performance improvements in the project came from moving away from a strictly one-term-at-a-time multiply-accumulate approach.

Measured end-to-end inference latency:

  • sequential-only version: 13,775,599 cycles
  • version with mac4_unit: 3,639,471 cycles

That is a 73.6% reduction in cycle count.

Latency analysis also showed that the early layers dominate runtime, especially conv1, because those layers still operate on the largest spatial dimensions. Later layers contribute less time as the feature maps become smaller.

Verification Strategy

The project includes both unit-level and end-to-end testbenches:

  • SourceFiles/tb_conv_layer.sv
  • SourceFiles/tb_maxpool_layer.sv
  • SourceFiles/tb_fc_layer.sv
  • SourceFiles/tb_alexnet_top.sv

These tests are intended to verify both local module behavior and full-pipeline sequencing.

The layer-level testbenches use small, hand-crafted inputs that are easy to inspect manually. The top-level testbench initializes the image and weights directly inside the design and checks that the overall inference path completes and produces the expected class.

Limitations

This project has several important limitations.

Not a full AlexNet implementation

This design does not match the full size of the original AlexNet. It uses reduced channel counts and only 10 output logits rather than 1000 ImageNet classes.

Not trained

This network was not trained on a real dataset, and the repository does not include trained AlexNet weights. The top-level testbench uses synthetic weights chosen for deterministic functional testing rather than learned parameters from training.

That means this project should be understood as a hardware architecture and inference-flow demonstration, not as a meaningful image classifier.

No real model accuracy claim

Because the network is not trained and does not use real learned parameters, it does not make sense to discuss classification accuracy in the usual machine-learning sense. The output is useful for validating dataflow and control, not for benchmarking recognition quality.

Simulation-oriented memory model

The design stores images, weights, and intermediate feature maps in internal arrays. This is convenient for simulation, but it is not the same as integrating with realistic on-chip or off-chip memory interfaces.

Testbench reaches into internal storage

Because weights and images are stored internally, the testbench writes directly into the top-level module arrays. This is convenient for testing but would not be the ideal interface style for a more polished accelerator design.

Limited optimization depth

The design parallelizes four terms at a time, which was a good compromise between simplicity and speed. More aggressive parallelization was possible, but it was not pursued further in order to keep the project manageable.

Future Improvements

There are several natural extensions that would make this design more complete:

  • add a real memory interface
  • reuse intermediate feature-map storage
  • increase datapath parallelism beyond four MAC terms
  • support loading trained weights from files or memory
  • expand the output layer and parameter set toward a larger AlexNet-like model
  • evaluate synthesis or FPGA deployment constraints rather than only simulation behavior

Conclusion

This project demonstrates a complete AlexNet-style inference pipeline in SystemVerilog using a design that is small enough to simulate and debug while still preserving the essential structure of the architecture.

The final system shows the tradeoff between simplicity and performance clearly: a mostly sequential FSM-based design is straightforward to understand, and a targeted datapath optimization like mac4_unit can still provide a large latency improvement without overwhelming the design with complexity.

Its main value is as a hardware CNN implementation and architecture study rather than as a trained classifier.

References

  • Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton, "ImageNet Classification with Deep Convolutional Neural Networks," NeurIPS 2012
  • AlexNet overview material from Wikipedia
  • SystemVerilog reference material from Accellera

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages