AI Product & Transformation Lead
Financial Services | Fintech | Applied GenAI
AI Product and Transformation Lead with more than 10 years across financial services, fintech, quantitative markets and digital platforms, currently Head of AI at a Swiss corporate finance advisory firm. Translates commercial and financial-services problems into prioritised AI use cases, then owns product vision, MVP scope, roadmaps, requirements, evaluation criteria and governance through to adoption. The current portfolio spans an ultra-low-latency AI inference cloud (Qorinix), a read-only AI money assistant over Open Banking (LaSpend), a Swiss consumer marketplace (Fixxmi) and an AI-native trading desk. Bridges executives, users, domain experts, technical contributors and vendors, and accelerates discovery and delivery through AI-assisted prototyping and coding agents. Comfortable deep in the technical detail, from LLM, RAG and agentic patterns to inference latency and cost trade-offs, evaluation harnesses and cloud delivery: able to shape requirements, evaluate outputs and direct governed delivery with confidence. MSc Computer Science, AI and Data Science (Merit).
Weidmann & Cie. AG
London, UK (Hybrid)December 2024 - Present
Swiss corporate finance advisory firm (M&A and growth advisory) building an applied-AI product portfolio: Qorinix (ultra-low-latency AI inference cloud), LaSpend (read-only AI money assistant over Open Banking), Fixxmi (Swiss consumer marketplace), pricing and insurance comparison workflows, and an AI-native trading desk.
Pacific Cloud Computing Ltd.
Hong Kong & Remote UKDecember 2021 - December 2024
Promoted internally from Senior Software Engineer to lead the firm's AI transformation across SaaS, analytics and market-intelligence products.
Pacific Cloud Computing Ltd.
Hong KongJanuary 2015 - November 2021
Groupon.com
Hong KongApril 2013 - December 2014
SoManyCall Telecom
Hong KongMarch 2008 - April 2013
High-speed AI inference cloud for latency-sensitive workloads: real-time agents, voice assistants, trading alerts and conversational AI. Targets sub-200ms TTFT (p50), 100-200+ tokens/sec and sub-10ms cache-hit latency at a fraction of frontier-API cost; internal benchmark targets place it 4-12x faster than leading low-cost providers and 6-12x faster than frontier APIs. Multi-tier model router with provider failover, streaming SSE, per-call cost ledger, usage-based billing and entitlements. Stack: Python, FastAPI, PostgreSQL, Redis, LLM Router, Stripe entitlements, OpenTelemetry.
Side-by-side real-time benchmark harness across 12 frontier and fast-inference providers. Streaming token-by-token with live leaderboard, TTFT p50/p95 analytics, tokens/sec, estimated cost-per-task and service-tier selectors. Per-provider model picker, run-count (1-20) for statistical stability and quick TTFT-turbo profiles. Next.js 14, TypeScript, server-side API isolation, streaming-SSE.
Read-only AI money assistant over regulated UK Open Banking: subscription detection, recurring-spend dashboard, waste scoring and assisted cancellation with verified action receipts. Never moves customer money; human-in-the-loop for all user-triggered actions; PII-minimised inputs and deterministic fallback when models fail. Cloudflare Pages + Functions + D1 architecture, React 19 + Vite + Tailwind frontend.
Swiss-first B2C service-job lead marketplace with AI-powered intent classification and lead matching across 12+ categories in DE-CH / EN. Pay-per-lead monetisation with Stripe Credit Packs, nFADP / GDPR-compliant data boundaries, Firestore-backed real-time admin dashboard. Next.js 15 static export + Firebase Cloud Functions (Node 20) + Firestore europe-west6.
Closed-loop flow applied across every firm AI product: policy check, evidence-pack retrieval, prompt control, model router, action runtime, audit ledger, policy feedback. Every call versioned, costed, support-traceable and replayable. Content-addressed evidence packs, per-product cost attribution, human-in-the-loop escalation for high-risk cases.
Distributed ML inference pipeline serving hundreds of predictions per day at <200ms TTFT across classification, sentiment, affordability scoring and anomaly detection. Shard-aware routing, warm-pool autoscaling, ONNX-compiled models with fp16 quantisation, cache-through Redis tiering, multi-region active-active with health-weighted failover. FastAPI, Redis, TimescaleDB, Triton.
Novel architecture fusing LLM narrative understanding with real-time HFT. Custom transformer with attention heads over microstructure features, ensemble gating and latency-aware inference scheduling that down-routes heavy heads on tight time budgets. AI-assisted development workflow with a three-gate validation framework delivered an 85% reduction in strategy development time. Python, PyTorch, CCXT, Ray Serve, Triton.
University of Wolverhampton, UK
2023 - 2025 | Grade: Merit
Hong Kong University of Science & Technology
1995
Open to AI Product, AI Transformation and Head of AI roles across financial services, fintech and applied GenAI.
Full UK right to work | Based near London (Reading) | Hybrid or remote | Languages: English, Cantonese, Mandarin
Last Updated: July 2026 | Portfolio: www.Philip.pm | References available upon request