> INIT: hpc.consulting --mode=maas --gpu=cuda

Stop bleeding compute capital.
Optimize your large algorithmic problems.

I architect Model-as-a-Service infrastructure and optimize low-level GPU workloads to slash your cloud bills. Fewer nodes. Higher throughput. Kernels that actually saturate the hardware you're already paying for.

///01 / the problem

Your GPUs are idle 80% of the time. You're paying full price for all of it.

Inefficient VRAM allocation, uncoalesced memory access, and Python-first wrappers turn a $40k/month compute budget into a runway-burning liability. Most teams scale horizontally — more instances, more nodes — instead of fixing what's already running.

  • ×Kernels launching without saturating the SMs.
  • ×Memory transfers hiding 3× more latency than the math.
  • ×LLM serving stacks reserving VRAM they'll never touch.
  • ×Autoscalers that scale spend, not throughput.
///02 / service catalog

Three engagements. One outcome: fewer dollars per FLOP.

S_01// service

MaaS Architecture Design

Architecting LLM workflow orchestrators, inference gateways, batching layers, and multi-tenant model routing built to scale without doubling your bill.

  • vLLM / TGI / Triton
  • Request batching
  • KV-cache strategy
  • Multi-tenant isolation
S_02// service

GPU Compute Profiling

Nsight Compute / Systems deep dives. I find the bottlenecks — memory coalescing, warp divergence, occupancy stalls — and ship a report with prioritized fixes.

  • Nsight Compute
  • Roofline analysis
  • Memory coalescing
  • Kernel timeline
S_03// service

CUDA / C++ Kernel Optimization

Replacing inefficient high-level wrappers with native performance. Custom CUDA kernels, C++ hot paths, and multi-GPU communication that actually uses NVLink.

  • Custom CUDA kernels
  • Tensor core usage
  • NCCL / NVLink
  • Fused ops
///03 / authority

$ whoami

role     : Independent HPC / GPU Engineer
research : M.Sc., HPC & Algorithms
lab      : University of Bergen (UiB)
focus    : Parallel algorithms
           Multi-GPU supercomputers
           Model-as-a-Service

Research-grade parallelism.
Production-grade delivery.

M.Sc. Researcher in HPC & Algorithms at the University of Bergen. Specializing in parallelizing algorithms for multi-GPU supercomputer environments — the same primitives that decide whether your inference cluster returns tokens in 40ms or 400ms.

I work with a small number of technical founders and CTOs per quarter. Every engagement starts with a fixed-scope Compute Audit — no retainers, no vendor lock-in.

// 04_contact

Ship faster kernels.
Spend less on GPUs.

Fixed scope, two-week Compute Audit. Deliverable: prioritized report with expected $/throughput impact per fix.

▸ Request a Compute Auditresponse < 48h · limited slots per quarter