Відомості ItsMeAjayKV
Pushing consumer hardware to its limits with local LLMs & OSS
I’m a software engineer with a background in AI and machine learning, exploring how much we can get out of local LLMs on hardware ordinary people can own. My setup is “poor man’s hardware”: a used RTX 3060 and a used RTX 3090.
I benchmark and profile models, experiment with quantization and speculative decoding, build evals, find what breaks, and fix what I can. More recently, I’ve been working on custom CUDA kernels and porting Escha W2 models to llama.cpp, getting deeper into the code that determines how efficiently models run. (https://github.com/Ajay9o9/llama.cpp-escha)
I also publish GGUF conversions and quantized models on Hugging Face, where my releases have received more than 15,000 cumulative downloads.
Hugging Face:
Through projects like **runs-on-a-3060** (https://github.com/Ajay9o9/runs-on-a-3060), I share practical recipes, commands, benchmarks, and instructions so others can reproduce the results on their own machines.
I believe knowledge and access to AI shouldn’t be controlled by a handful of powerful companies. Local LLMs and open source give people more control over the tools they use, and there’s still plenty of work to do to make capable models run well on affordable hardware.
I share my work openly on GitHub and Hugging Face, and post experiments, results, and development updates on X (https://x.com/ItsmeAjayKV) including the things that don’t work.
My goal is to keep contributing useful work to the local LLM community while going deeper into inference optimization and LLM research. If you find my work useful, your support helps cover compute costs and gives me more room to experiment, build, and share what I learn.
Нещодавні прихильники

