Mathematical AI Researcher · Ph.D. Candidate · University of Central Florida

Kasun Dewage

I work on the mathematics underlying artificial intelligence — tensor decomposition, low-rank structure, random matrix theory and numerical linear algebra — to understand how large transformer models represent learned information and to make them more efficient to adapt and compress.

Tensor methods Parameter-efficient fine-tuning Attention structure Random matrix theory Quantitative finance
Portrait of Kasun Dewage

Hover to rebuild the mesh

6

First-authored papers
accepted in 2026

840×

Fewer adaptation parameters
than LoRA on LLaMA3-8B

4.76/5

Instructor effectiveness,
Spring 2026 (dept. 4.21)

2.01M

Minute-level observations
in the forecasting study

Research focus

Mathematics for Efficient
Artificial Intelligence

I study the mathematics underlying artificial intelligence — tensor decomposition, low-rank structure, random matrix theory, and numerical linear algebra — to understand how large transformer models organize learned information and why they can often be adapted in surprisingly low-dimensional spaces.

I use this structure to develop methods for adapting, compressing, and understanding large AI models with far fewer trainable parameters and lower computational cost. A central focus is CRAFT, a parameter-efficient fine-tuning framework that exploits shared structure across transformer layers, with the broader goal of making advanced AI more efficient, affordable, and accessible.

I also develop systematic trading models and equity-market simulations using machine learning, time-series modeling, and quantitative methods.

Tucker tensor decomposition above a point-cloud progression of human evolution 𝒳 ≈ 𝒢 ×₁U⁽¹⁾ ×₂U⁽²⁾ ×₃U⁽³⁾

Direction 01 · Adaptation

Parameter-efficient fine-tuning

Decompose the pre-trained weights, freeze the factors, train only what is small.

  • Cross-layer Tucker-3 / HOSVD on stacked attention weights
  • Adaptation cost independent of model width and depth
  • CRAFT, main-conference paper at COLM 2026

Direction 02 · Structure

What attention actually stores

Spectral and statistical analysis of pre-trained projections across 11 models.

  • Marchenko–Pastur separation of signal from bulk
  • Calibration-free structured head pruning
  • Component type predicts quantization sensitivity

Direction 03 · Markets

Foundation models on price data

Frozen time-series foundation models, corrected by small neural and classical heads.

  • Frozen 200M-parameter TimesFM backbone
  • Mean per-day correlation 0.059 → 0.373
  • Accepted at IEEE IJCNN / WCCI 2026

Headline result · COLM 2026

CRAFT matches LoRA
on a rounding error
of its parameters.

LoRA-CRAFT applies Tucker decomposition to attention weights stacked across layers, freezes every resulting factor, and trains only small square transformations on those factors. On LLaMA3-8B it beats the reported LoRA commonsense average while training roughly one eight-hundred-fortieth as many adaptation parameters.

Adaptation parameters LLaMA3-8B
LoRA 57 M
CRAFT 0.068 M

0.12% of the bar above — drawn to scale.

Commonsense accuracy, average %
LoRA 80.8
CRAFT 82.9

Selected papers

COLM 2026 Conference on Language Modeling Main conference

LoRA-CRAFT: Cross-layer Rank Adaptation via Frozen Tucker Decomposition of Pre-trained Attention Weights

Kasun Dewage, Marianna Pensky, Shankhadeep Mondal, Suranadi De Silva

An extremely parameter-efficient fine-tuning method that performs full Tucker decomposition by higher-order SVD directly on pre-trained weights organised as cross-layer 3D tensors, freezes all resulting factors, and adapts the model through lightweight trainable transformations on each factor matrix.

IJCNN 2026 IEEE WCCI Accepted

Hybrid Neural-Classical Correction for Frozen Time Series Foundation Models

Kasun Dewage, Suranadi De Silva, Shankhadeep Mondal

A comprehensive ablation on high-frequency stock prediction: a frozen 200M-parameter TimesFM backbone paired with small learned correctors, evaluated over 2,011,399 one-minute observations across ten technology stocks.

ICMLA 2026 IEEE Accepted

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention

Kasun Dewage, Marianna Pensky, Suranadi De Silva, T. H. Bandara

Marchenko–Pastur random matrix theory separates each attention projection into a random-like bulk and a set of spectral outliers. Zeroing the outliers in Mistral-7B drives HellaSwag, MMLU and PIQA close to chance; zeroing a count-matched bulk subset does far less damage.


Talks & scholarly presentations

Research presented across
AI and computational intelligence.

Conference presentations of my work on efficient foundation models, tensor methods, transformer structure and model compression.

IEEE WCCI / IJCNN 2026 Maastricht, the Netherlands · June 21–26, 2026 Presented

Hybrid Neural-Classical Correction for Frozen Time Series Foundation Models

IEEE World Congress on Computational Intelligence · International Joint Conference on Neural Networks

A presentation on correcting a frozen 200M-parameter time-series foundation model with lightweight neural and classical components for high-frequency financial forecasting.

IEEE ICNLP 2026 Xi’an, China · March 20–22, 2026 Presented

On the Compressibility of Fine-Tuned Attention: A Tensor Decomposition Study of PEFT Weight Structure

International Conference on Natural Language Processing

A tensor-decomposition study of the structure and compressibility that emerges in parameter-efficiently fine-tuned transformer attention weights.

COLM 2026 San Francisco, California, USA · October 6–9, 2026 Expected

LoRA-CRAFT: Cross-layer Rank Adaptation via Frozen Tucker Decomposition of Pre-trained Attention Weights

Conference on Language Modeling · Main conference

Expected scholarly presentation of CRAFT, a cross-layer Tucker/HOSVD approach to extremely parameter-efficient adaptation of transformer attention.

IEEE ICMLA 2026 Rochester, Michigan, USA · October 5–7, 2026 Expected

Accepted work on transformer compression, quantization sensitivity and spectral structure

International Conference on Machine Learning and Applications

Expected presentations of accepted ICMLA 2026 work on magnitude-profile attention-head pruning, component-level quantization sensitivity, and spectral outliers in transformer attention.


Teaching

Six years as
instructor of record.

Calculus II, Calculus III and Linear Algebra at UCF, and a full service load at Sam Houston State before that. Student evaluations have stayed between 4.6 and 5.0 out of 5.0, consistently above both department and college averages.

Spring 2026 · MAC2312 · Instructor effectiveness 4.76 / 5.0
Kasun Dewage Department

Made clear efforts to engage students4.88

Helpful in responding to questions4.81

Enhanced my understanding of the material4.69


Education

Mathematics, computer vision,
and quantitative training.

Graduate training spanning mathematical foundations, computer vision, and quantitative finance.

Expected Fall 2026

Ph.D. in Mathematics

University of Central Florida · GPA 3.7

Advisor: Prof. Marianna Pensky

2024

M.S. in Computer Vision

University of Central Florida · GPA 3.8

UCF reported its computer vision research as No. 8 nationally in 2025.

2017

M.S. in Pure Mathematics

Sam Houston State University

2014

M.S. in Financial Mathematics

University of Colombo

2012

B.Sc. in Physical Science

University of Colombo


Contact

Open to postdoctoral, faculty and research roles from Fall 2026.