Welcome — I build things that matter
Parsa Rostamzadeh
Graduate Research Assistant
Building intelligent systems at the intersection of software, hardware, and machine learning.
Get to know me
About Me

Graduate Research Assistant · Paderborn University
I'm a computer engineer and research assistant at Paderborn University, focusing on approximate computing, hardware-aware machine learning, and FPGA-based neural network optimization. My research centers on making deep learning deployable on resource-constrained hardware — without sacrificing more accuracy than necessary.
I build end-to-end pipelines that span the full stack: from quantization-aware training and circuit synthesis to multi-objective design space exploration. When I'm not optimizing circuits, I work on graph neural networks and explainability — understanding not just what models predict, but why.
What I Do
I build approximate computing pipelines and hardware-aware ML systems — from FPGA-deployed neural networks to cross-layer circuit synthesis.
What Drives Me
Making neural networks smaller and faster without breaking them. Approximate computing lets me trade a little accuracy for a lot of efficiency — and I find that trade fascinating.
Current Focus
Cross-layer approximate synthesis for FPGA-deployed neural networks — profiling sensitivity, generating approximate neuron variants, and exploring Pareto-optimal area-accuracy trade-offs.
What I work with
Skills & Tech Stack
ML / AI
Hardware
Synthesis and verification at the gate-level, optimizing for specific silicon constraints.
- FPGA
- RTL Design
- Circuit Design
- Standard Cell Design (VLSI)
- Xilinx Vivado
- Yosys / ABC
- LSOracle
- Icarus Verilog
- BLASYS
Languages
Tools & Infra
What I've built
Projects
CLAS — Cross-Layer Approximate Synthesis
End-to-end pipeline for approximating FPGA-deployed neural networks at the circuit level.
A four-stage research pipeline that takes a quantized neural network, profiles per-neuron sensitivity via LASSO regression, generates approximate neuron variants using BLASYS, and runs a multi-objective genetic algorithm to find Pareto-optimal area-accuracy trade-offs across the full network. Evaluated using Vivado synthesis for area and Icarus Verilog for accuracy.
BLASYS Neural Network Approximation Pipeline
HPC-optimized pipeline for approximating neural network neurons using Boolean Matrix Factorization.
Automates the approximation of neural network neurons using Boolean Logic Approximation and Synthesis (BLASYS). Extracts neurons from network layers, applies Boolean Matrix Factorization with configurable error thresholds, and generates a Pareto front of approximate designs trading off chip area vs. accuracy. Optimized for 120-core HPC nodes with full parallel execution.
Circuit Approximation via Spectral Partitioning
Selectively approximates arithmetic circuits to minimize silicon area within a user-defined error budget.
A three-stage pipeline that analyzes circuit sensitivity via reverse-topological DFG traversal, partitions circuits using the Fiedler vector of the graph Laplacian with sensitivity-aware edge weights, then applies Lagrangian relaxation to optimally distribute an error budget across partitions. Achieves near-global-optimal approximation with independent subproblem solving.
Knowledge Graph Validation with SHACL & Ontologies
Team project validating knowledge graphs against SHACL constraints in the presence of a DL-Lite_R ontology.
A six-student project group at Paderborn University implementing Ahmetaj et al. (ECAI 2023): validating knowledge graphs against SHACL shapes while respecting what a DL-Lite_R ontology entails. I worked on the materialisation approach, building an austere canonical model (a bounded "chase") so plain pySHACL sees entailed facts, plus the validation pipeline, stratification pre-check, stress tests, interactive visualisation, and performance profiling.
PubMed Graph Attention Network + XAI
Graph Attention Network for scientific paper classification with full explainability framework.
Implements a GAT for node classification on the PubMed citation network (diabetes literature), featuring a comprehensive Explainable AI framework with attention pattern analysis, feature importance visualization, and multi-perspective explanations. Built at Paderborn University for the Explainable AI course.
Reservoir Network Quantization
Quantization experiments on Echo State Networks for NARMA time series prediction.
Investigates the effect of weight quantization (4-bit, 6-bit, 8-bit) on Echo State Network performance across NARMA10 and NARMA20 tasks. Includes quantized weights for reservoir, input, bias, and readout layers, with per-level accuracy comparisons and full saved states for reproducibility.
Neural Network from Scratch
Full feedforward neural network built with NumPy only — no frameworks.
Implements every component of a neural network by hand: activation functions (Sigmoid, ReLU, Tanh) with derivatives, Binary Cross-Entropy and MSE loss, Xavier and He weight initialization, forward propagation, backpropagation via chain rule, and gradient descent. Applied to breast cancer malignancy classification on real clinical data.
CIRCA — Approximate Circuit Generation
Extensions and implementations on the CIRCA approximate circuit synthesis framework.
Contributed implementations and extensions to CIRCA, the modular approximate circuit generation framework by Paderborn University. Work includes evolutionary approximate circuit variants (Circa_evo) and a DDECS-targeted extension integrating custom approximation strategies into the CIRCA pipeline.
My journey
Experience & Education
Graduate Research Assistant
Computer Engineering Group, Paderborn University
- ▸Co-developed CLAS, a cross-layer approximate synthesis framework for LUT-based DNN accelerators (ARC 2026).
- ▸Co-developed a partition-based design-space exploration framework for approximate accelerators (under review).
- ▸Extended CIRCA, the group’s approximate circuit generation framework, with evolutionary approximation variants.
- ▸Automated parallel synthesis and simulation campaigns on the Noctua 2 HPC cluster (SLURM).
M.Sc. in Computer Engineering
Paderborn University
- ▸Studying for a Master’s in Computer Engineering, specializing in Embedded Systems.
- ▸Worked on many courses and projects related to hardware (FPGA, ASIC).
- ▸Completed in-depth projects in LLMs and XAI, including LLM fine-tuning and RAG pipelines.
Software Developer
Hesab Rayan Pars
- ▸Gained experience working in a team on C# accounting software development.
- ▸Got hands-on experience with real-world software development and architecture.
B.Sc. in Computer Engineering
Tehran Azad University
- ▸Built a strong programming foundation, especially in C#, C++ and Python.
- ▸Completed many implementations and personal projects in web development with modern frameworks (.NET, FastAPI).
- ▸Gained a solid understanding of computer architecture and organization.
Research output
Publications
Divide et Approxima: Scalable Design Space Exploration for Approximate Accelerators via Partitioning and Sensitivity-driven Error Allocation
Submitted to Design, Automation and Test in Europe Conference (DATE) 2027 · awaiting acceptance
Frameworks for automated synthesis of approximate hardware accelerators rely on repeated, typically simulation-based error estimation to navigate a design space that grows exponentially with the number of approximable components. For larger accelerators, this estimation cost dominates and prohibits a single exploration run from effectively exploring the solution space. We address this challenge with a partition-based synthesis framework that decomposes an accelerator into sub-designs small enough to explore tractably, combined with a lightweight, sensitivity-driven error-budget allocation strategy that distributes the global error across partitions so they can be searched independently and in parallel. Across four accelerator kernels in 20 size configurations, spanning a 28× size range and error targets of 1–10% RMSE, the framework accelerates design space exploration by up to 92.9× while meeting every error constraint. The resulting designs are never Pareto-dominated by single-search design space exploration and achieve up to 24.6 pp higher area savings.
CLAS: A Cross-Layer Approximate Synthesis Framework for LUT-based DNN Accelerators
A. Jafari, A. H. Hadipour, M. Awais, M. Rostamzadeh-Khameneh, H. Ghasemzadeh Mohammadi, M. Platzner
22nd International Symposium on Applied Reconfigurable Computing (ARC) · Springer LNCS · Cagliari, Italy
LUT-based DNN accelerators offer ultra-low latency FPGA inference, but their adoption is severely constrained by excessive resource consumption. This paper introduces CLAS, a cross-layer approximation framework for LUT-based DNNs that redefines approximation in fully unrolled networks by treating neurons as substitutable RTL components and jointly exploring combinations of approximated layers at the RTL across the network. CLAS achieves substantial LUT reductions with a small accuracy loss, outperforming algorithmic-level approximation baselines by delivering an additional 33% area savings with only a 4% drop in classification accuracy on the MNIST dataset.
Get in touch
Let's Work
Together
Open to research collaborations, full-time roles, and interesting side projects. Feel free to reach out.