Software engineer. Computer Science and Engineering at BUET.
I build systems where a language model has to be right, not just fluent.
Website · Blog · LinkedIn · Devpost · Email
Most of my work sits where an AI system meets a requirement to be correct. The principle underneath all of it is the same: give the model the smallest job that only it can do, and make everything around it deterministic, measurable and reproducible.
I also run SrotDev, a flat engineering collective I founded so the team ships under one name instead of scattering work across individual accounts.
Thermal imaging, clinical speech, medical documents, loan covenants. The wins have almost nothing in common, which is the part I am proudest of.
| Award | Competition | Project |
|---|---|---|
| Grand Prize Winner | FLIR App Challenge 2025 to 2026 | SolarSnap |
| Winner, Best Overall Project | ML Empowerment Build Challenge 2.0 | CADENCE |
| Winner, Best ERNIE Multimodal Application | ERNIE AI Developer Challenge, Baidu | Doclyst |
| Winner, Honourable Mention | LMA EDGE Hackathon, presented at the London finale | Coven |
| Project | What it does |
|---|---|
| HALFSPREAD | An options agent that prices its exit before it enters. Every published figure re-derives from an append-only journal with no key and no network. |
| Cassandra | Continuous integration for spreadsheets. An agent fleet finds the defect, writes the fix, and proves it by recalculating the workbook. |
| Cascade | Agent memory that knows when it has expired. Procedures are pinned to the policy versions they were derived from and refuse themselves when one moves. |
| DeepSIFT | 148 MCP forensic tools with per-claim grounding verification and a signable chain of custody. 4/4 against published ground truth, zero hallucinations. |
| CADENCE | Parkinson's screening from a 30 second voice sample, validated across three language corpora with cross-database testing. |
| Gotcha! | An AI that makes one deliberate mistake per challenge, so students learn to catch it. Generation and grading are separated so it cannot mark its own homework. |
More at ahammadshawki8.github.io/projects.
Where the scoreboard is a metric rather than a panel.
| Competition | Task and result |
|---|---|
| Olikobochon, IUTCS Datathon 2.0 | Decide whether a fluent Bengali answer is actually true. Third on the public leaderboard at 0.922 F1, using a deterministic decision ladder that sends only the residual to an LLM judge. Code |
| DL Sprint 4.0, BUET CSE Fest 2026 | Bangla speech recognition on hours-long audio with overlapping speakers, background music and long silences. Scored on word error rate weighted by sentence length. |
Languages Python · C/C++ · TypeScript · JavaScript · Java · SQL · Dart
Web React · Next.js · Django · FastAPI · Flask · Node.js · Tailwind · Flutter · JavaFX
AI and ML PyTorch · scikit-learn · MCP · Google ADK · RAG · Gemini · Claude · Pandas · NumPy
Infrastructure Docker · Google Cloud · AWS · PostgreSQL · MongoDB · Redis · GitHub Actions · Prometheus · Grafana
I publish long-form technical articles on my blog, and two of my pieces were published on freeCodeCamp News.
Open to internships, contract work and collaboration.
Get in touch





