FlashKernel
Custom CUDA C++ and Triton kernels for transformer inference — tiled FlashAttention, fused GeLU, RoPE, paged KV-cache — benchmarked with Nsight Compute on T4.
View details →Abdeljalil Jliouat Applied AI Scientist · Paris
I build efficient, reliable AI systems — from GPU kernels to robotic learning. Here I share the code, experiments and engineering decisions behind my work.
Three connected areas of applied research and engineering.
Efficient models, CUDA kernels and the systems that run them.
Learning from language, interaction and brain signals.
Classical and quantum methods, tested against practical constraints.
Explore the architecture, experiments and source code behind four systems.
Custom CUDA C++ and Triton kernels for transformer inference — tiled FlashAttention, fused GeLU, RoPE, paged KV-cache — benchmarked with Nsight Compute on T4.
View details →Language-grounded robotic manipulation — VLM planner decomposes instructions into sub-tasks, RL policies execute each step in MuJoCo simulation.
View details →Foundation model for neural signal decoding — pre-train a transformer on large-scale EEG, fine-tune for motor imagery BCI with frequency-band attention.
View details →Quantum-classical hybrid optimization for energy grids — QAOA and VQE applied to unit commitment on real ENTSO-E data, benchmarked against MILP solvers.
View details →Leading applied ML and AI systems in startups, NGOs, and banking.
Available for collaborations, advisory work, and technical leadership roles.