A Socially-Aware, Agentic AI & Federated AI Hardware Accelerator designed in SystemVerilog and verified with Cocotb.
The AAA core is a custom RTL (Register Transfer Level) design tailored for edge AI workloads. Moving beyond standard matrix multiplication, this architecture integrates Federated Learning aggregation natively in hardware and employs an Agentic Guard to ensure physical execution security against invalid software modes.
This project demonstrates a complete Hardware-Software Co-Design flow, verified using a 100% Python-based testbench.
The mode_request bus is 3 bits wide, allowing for 8 distinct hardware states across the 4 parallel Processing Elements (PEs):
- Mode 0 (
000): NOP (No Operation) / System Idle. - Mode 1 (
001): Deep Learning MAC with strict overflow saturation. - Mode 2 (
010): Machine Learning Distance Search (KNN). - Mode 3 (
011): Liquid Neural Networks (Euler ODE solver). - Mode 4 (
100): Reinforcement Learning (Hard-bounded saturation). - Modes 5, 6, 7 (
101-111): Unmapped / Forbidden execution states.
Instead of relying on software to merge AI models, the hardware natively combines intelligence. It computes a 50/50 weighted average of local device learning and global cloud weights using arithmetic right shifts to preserve signed integer logic (negative weights) with single-cycle efficiency:
A hardware-level Finite State Machine (FSM) that isolates the data path and actively monitors the mode_request bus. If software attempts to execute the unmapped instructions (Modes 5, 6, or 7), the Limiter acts as a "pain reflex." It intercepts the clock gating, prevents PE execution, and immediately flags an error to the central system via bit 2 of the system_status bus.
The AAA_Top module is parameterized for N compute lanes (default N=4).
| Port Name | Direction | Width | Description |
|---|---|---|---|
clk |
Input | 1-bit | Master system clock |
reset_n |
Input | 1-bit | Active-low asynchronous reset |
task_trigger |
Input | 1-bit | Initiates the Scheduler FSM |
mode_request |
Input | 3-bit | Software request for the 8-state PE execution mode |
virtual_data_bus |
Input | 64-bit | Packed input data ( |
virtual_weight_bus |
Input | 64-bit | Packed local weights ( |
global_weight_in |
Input | 16-bit | Broadcast weight from the cloud for aggregation |
system_status |
Output | 3-bit | Status Flags: [2] ERROR/REJECT, [1] IDLE, [0] BUSY
|
result_stream |
Output | 128-bit | Packed 32-bit accumulators from all 4 PEs |
- Hardware Description Language: SystemVerilog (IEEE 1364-2005)
- Software Control/Verification: Python 3.14 (Cocotb v2.0.1)
- Simulator: Icarus Verilog v12.0
- Development Environment: MSYS2 / UCRT64 on Windows 11
AAA_Master_Project/
├── design.sv # RTL Top-level, SIMD PEs, Aggregator, and Scheduler
├── test_aaa.py # Cocotb Python verification testbench & runner API
└── Makefile # Legacy build instructions (Bypassed via Python Runner)
The Python testbench (test_aaa.py) operates as the software agent, injecting data into the virtual silicon to verify:
[SILICON PROOF] Federated Result: 10000000000000000000000000000000100000000000000000000000000000001000000000000000000000000000000010000000000000000000000000000000
[AGENTIC CHECK] Requesting Forbidden Mode 7... [SILICON PROOF] SUCCESS: Hardware Limiter Triggered & Mode 7 Rejected!
** TEST STATUS SIM ** test_aaa.aaa_verification PASS
- Aggregator Accuracy: Confirms the hardware correctly scales and merges local weights (e.g., 100) with global weights (e.g., 50) before executing the MAC operation.
- Security Integrity: Intentionally requests an invalid state (
Mode 7) to verify the hardware's internal Limiter triggers theREJECTstate and protects the compute lanes.
Ensure you have Icarus Verilog and Python installed. For Windows users, an MSYS2 environment is recommended.
- Clone this repository:
git clone [https://github.com/YourUsername/Autonomous-Adaptive-Accelerator-AAA.git](https://github.com/YourUsername/Autonomous-Adaptive-Accelerator-AAA.git) cd Autonomous-Adaptive-Accelerator-AAA
2.Install the Verification Framework:
pip install cocotb
3.Run the python runner script to compile and test silicon
python test_aaa.py
4.Look for the green PASS output in the terminal, confirming both the Federated Result and the Hardware Limiter trigger.