Back to Groups
Independent Research
Research in efficient LLM inference, quantisation algorithms, and hardware-aware machine learning.
R
Designed & implemented Variable Bit Floating-Point (vBFP) quantisation on Llama-3-8B.
C++CUDAPyTorchPython
View Case Studyl
An ambitious project involving the deployment and serving of large language models.
PythonPyTorchLLMsDocker
View Case Study