Back to Groups

Independent Research

Research in efficient LLM inference, quantisation algorithms, and hardware-aware machine learning.

R

Row-Tiered Mixed-Precision Quantisation for LLMs

Designed & implemented Variable Bit Floating-Point (vBFP) quantisation on Llama-3-8B.

C++CUDAPyTorchPython
View Case Study
l

llm.karanmertiya.me

An ambitious project involving the deployment and serving of large language models.

PythonPyTorchLLMsDocker
View Case Study