- SDE intern at Clickpe (YC W23), Dec 2025 to present
- GSoC contributor, June 2025 to September 2025
01Deep learningdocs/dl/
-
The 32-Bit Trap: how a 4 GB tensor becomes a wild pointer
Fixing the 32-bit addressing trap of CuTe DSL encountered while writing qwen2.5 14b megakernel
-
Vectorization and coalescing: the two ways to move memory fast
Two independent knobs decide kernel bandwidth: how wide each thread's load is, and whether adjacent lanes touch adjacent addresses. CuTe gives you exactly one call for each.
-
Optimizer state: the memory bank of deep learning
What AdamW remembers per parameter, why bias correction exists, and why mixed precision forces an FP32 master copy of the weights.
-
Visualizing LLM tensors: from tokens to batch matrices
From tokenizer IDs to padded batches and attention masks: every shape a prompt passes through, drawn so every number fits on screen.
-
PyTorch basics: tensors, autograd, and the training loop
Tensor mechanics, autograd with a hand-checked gradient example, the five-line training loop, and the pitfalls behind ninety percent of runtime errors.
02Systems and infrastructuredocs/
-
Deploying agents on scale: what nobody tells you about scaling coding sandboxes
Queues instead of open connections, ephemeral sandbox pods instead of bare shells, JSONL logs instead of memory pressure. What breaks when ten agents share one VM.
-
Building a tiny pub/sub system in TypeScript
One Map, one Set, async iterators: a GraphQL-ready event bus in a single dependency-free class.
03Open sourcedocs/
-
How to start contributing to open source: a guide for newcomers
Tiny doc fixes this week, Hacktoberfest in October, GSoC or LFX within a year: contribution types, program calendars, skills, and a repeatable routine.
04Elsewhere
- Email: manishbiswal754@gmail.com
- GitHub: github.com/iamanishx
- LinkedIn: linkedin.com/in/manish-biswal-xd
- Twitter / X: x.com/iamanishx
- Resume: mera_resume.pdf