Founder of TrackHRS · live on the market

I design systems that scale.

I'm Ahmed Ali, a software architect and engineering lead. I design and build event-driven backends, web and ecommerce platforms, cloud infrastructure and production AI systems, and lead the teams that ship them.

Open to remote and hybrid work worldwide
ahmed@prod: activity-pipeline
$ kafka-consumer --topic activity --group activity-workersconsumer group balanced · 3 partitions · lag 0✓ redis dedup      12,481 events   0 duplicates✓ batch writer     100 items / 5s window✓ bullmq classify  queued 312 · failed 0✓ cache invalidate dashboards freshp99 write latency  38ms$ 

500+

Projects delivered

Web, ecommerce, CMS, ERP and AI platforms

6+

Years engineering

FinTech, AI and SaaS at production scale

100K+

Events per day

Kafka pipelines running in production

20+

Engineers led

Across AI, ecommerce and CMS product lines

What I do

Four things I build, end to end

From the product your customers use down to the pipelines and infrastructure underneath it, designed, built and taken to production.

Product & platform engineering

The things your customers actually touch, built to hold up once real traffic arrives.

  • Web applications & dashboards
  • Marketing sites & landing pages
  • Ecommerce platforms
  • CMS & content platforms
  • ERP systems & integrations
  • Mobile & cross-platform desktop apps
  • MVPs taken from zero to launch

Backend & system architecture

The layer underneath: service boundaries, data flow and the trade-offs written down before the first commit.

  • System design & architecture reviews
  • Microservices & domain-driven design
  • Event-driven pipelines (Kafka, RabbitMQ)
  • REST & gRPC API engineering
  • Database strategy across SQL and NoSQL
  • Legacy modernisation & migration plans

AI, agents & automation

AI that ships as a service with failure paths, tracing and evaluation, not a demo that degrades quietly for six months.

  • Chatbots & conversational agents
  • LangGraph & multi-agent workflows
  • Custom MCP servers & tool integration
  • RAG pipelines & vector search
  • Agent SDK & generative APIs
  • LangSmith tracing & evaluation
  • Analytics & intelligence engines
  • Workflow & process automation

Cloud, scaling & audits

Making deployment boring, systems cheaper, and telling you honestly which part is actually the problem.

  • Kubernetes, Docker & CI/CD
  • Performance tuning & caching strategy
  • Cloud cost optimisation
  • Observability & production readiness
  • Architecture & code audits
  • Security & reliability reviews
Reference architectures

How I'd build what you're asking for

The four systems clients ask for most often, and the shape each one takes before a single line is written. Yours will differ in the details, but the structure rarely does.

AI assistant / RAG system

Answers from your own documents, with citations and a real answer when it does not know.

Source dataDocs, DBs, APIs
Chunk & embedStructure-aware
Vector storeIndexed + metadata
RetrieveTop-k + rerank
GenerateLLM + citations
EvaluateRecall + guardrails

Multi-agent & automation platform

Agents that complete a process by calling real tools, with every step reviewable rather than one opaque generation.

TriggerPrompt, event or schedule
PlanDecompose the task
Specialist agentsOne decision each
Tool callsMCP + your APIs
VerifyChecks before commit
Act & logWrite back + audit

Ecommerce & platform backend

Storefront, checkout and everything behind it, built so a traffic spike is a scaling event, not an outage.

StorefrontNext.js + cache
Catalog & searchIndexed reads
Cart & checkoutIdempotent writes
PaymentsGateway + webhooks
Order eventsQueue-backed
ERP / fulfilmentSynced + reconciled

Analytics & intelligence engine

Raw operational events turned into numbers people trust, with dashboards that stay fast as volume grows.

IngestEvents + integrations
DeduplicateIdempotency keys
Batch & storeBulk writes
EnrichClassify + score
AggregatePre-computed views
ServeCache-first APIs
Selected work

Systems I designed and shipped

Four platforms, each with a different constraint at its centre: throughput, precision, latency or ambiguity. Every case study walks the architecture and the full pipeline, not the feature list.

These four go deep. Behind them sit 500+ delivered projects, covering ecommerce and CMS platforms, ERP integrations, AI chatbots and agents, automation and internal tools, across client work and products of my own.

See all case studiesSee all case studies
How I think about systems

Every arrow in a diagram is a decision

Architecture is mostly about choosing what happens when something goes wrong. Three rules shape almost everything I build.

Decouple what fails differently

Ingestion, processing and delivery each break for their own reasons. An event backbone between them means one slow stage never stalls the rest.

Make correctness explicit

Deduplication keys, idempotency and distributed locks are mechanisms, not hopes. If two replicas can process the same event, something must decide which one counts.

Design the failure path first

Retries, circuit breakers and dead-letter queues get designed alongside the happy path, because the happy path is not the one that pages you at 3am.

What that has produced

60%+

Lower read latency

Redis cache-first strategy on a FinTech platform

25%

Cloud cost reduction

GCP rightsizing and observability work

~70%

Less deploy effort

Kubernetes and Docker replacing manual releases

50%+

Fewer integration issues

Clear service boundaries and ownership

Toolkit

The stack I reach for

Chosen per problem rather than per fashion. These are the tools I've taken to production often enough to know their failure modes.

NestJS
Go
TypeScript
Python
Rust
Apache Kafka
RabbitMQ
Redis
MongoDB
PostgreSQL
Kubernetes
Docker
FastAPI
Next.js
React
Spring
Google Cloud
Nginx
Jenkins
Tauri

System design · Backend & APIs · Messaging & streaming · AI & agent engineering · Databases · Cloud & DevOps · Frontend & desktop

See the full toolkit
Writing

Notes on building systems

Mostly the things I wish someone had told me before the incident, not after.

Have a system to build?

Whether it needs designing from scratch or rescuing from its own success, tell me what you’re building and I’ll tell you how I’d architect it.

Open to remote and hybrid work worldwide