AI ENGINEERING · AI ARCHITECTURE · LLMOPS · PLATFORM · SRE

Chris
Whyland

AI Engineer and Architect — agents, evals, and production LLM systems.

25 years of platform and SRE discipline applied to systems that think.

AI+
Focus
Agents · evals · LLMOps
DGX Spark
Private inference cluster
397B
Parameters
Largest model served
CVE
Django
CVE-2026-48588 credit

Security Research

Credited by the Django project for reporting CVE-2026-48588 (potential exposure of private data via cached Set-Cookie responses). Severity: low. Official listing: NVD / NIST.

Experience


Cognizant 2023 – Present
Senior Manager, DevOps Engineering
Fortune 500 consulting engagements across financial services, healthcare, and retail.
  • Architected SRE transformation for a Fortune 500 financial services client — full-stack maturity assessment, prioritized improvement backlog, and tailored training curriculum — improving availability, reducing recurring incidents, and increasing automation to reduce toil
  • Assisted a financial services client with design and implementation of agentic evaluation capabilities in their Snowflake Cortex Agent environment — evaluation harnesses and quality gates for enterprise agent workflows (client confidential)
  • Led GKE Kubernetes load testing and right-sizing for a major national retailer ($5B+ revenue) — validated 3× Black Friday capacity headroom through HPA tuning, cluster autoscaling, and pod resource optimization
  • Delivered SRE maturity assessment for a Fortune 50 healthcare company based on Google SRE practices — became the foundation of their reliability practice roadmap
AI R&D alongside professional role — 2024–Present
Jetty 2022 – 2023
Manager, DevOps Engineering
Insurtech startup. Led DevOps/SRE function — owned CI/CD, IaC, and the full AWS footprint.
  • Led containerization of legacy Python services to ECS Fargate — $800K annual cost reduction, 70% performance improvement
  • Ground-up IaC rewrite in Terraform with custom modules + GitHub CI (TFLint, TFSec, vulnerability scanning) — 30% reduction in deployment errors
  • Architected multi-account AWS infrastructure aligned to the Well-Architected Framework
  • Established SLO/SLA framework achieving 99.95% reliability using Datadog and Sentry
Covetrus 2020 – 2022
DevOps Manager / Lead DevOps Engineer
Fortune 1000 ($4.3B revenue) veterinary health technology. Recruited as Lead, promoted to Manager.
  • Unified 25-person global engineering team post-merger (Azure + AWS), creating governance for coordinated delivery across international markets
  • Microservices modernization: Terraform, Kubernetes, Kong, Confluent Kafka, GitLab CI, Harness — 60% reduction in time-to-market
  • SAFe Product Owner. Zero audit findings across SOX and PCI. $2M annual cloud cost reduction. GOAT Award Q2 2022.
EVO Payments International 2017 – 2020
Senior Lead Infrastructure Engineer / Enterprise Architect
Global payments platform operating across North America and Europe.
  • Established and led Systems Architect group supporting payment processing expansion into 5 new markets
  • DR strategy reducing potential revenue impact by 80%. SRE practice across 5 global dev centers — 60% MTTR reduction
  • Mentored 15+ senior engineers, building technical leadership pipeline that supported 3× team growth
Earlier Career 2002 – 2016

Progressive roles across DevOps engineering, systems administration, network engineering, and infrastructure — building the deep Linux, networking, virtualization, and storage foundations that underpin modern cloud and AI platform architecture. Includes DevOps Engineer at Vets First Choice/Covetrus, Systems Engineer at EVO Payments, Senior System Administrator at MEMIC, and earlier roles in consulting, technical support, and national-scale deployments.

Independent AI Infrastructure R&D


Self-directed research conducted alongside professional work, 2024–present. Private GPU hardware, production-pattern systems, real evaluations — not proofs of concept.

Private GPU Inference Cluster

2× NVIDIA DGX Spark · vLLM · Tensor Parallel · NVFP4

Dual-node GB10 cluster serving 30B–397B-class models, including 1M-context vision/reasoning. Currently tested and running: deepseek-v4-flash, deepseek-v4-flash-vision, glm-5.3-flash, and Qwen3.5-397B. Profile-switched serving with quantitative admission (KV / context), Prometheus/Grafana, and a hardware-specific vLLM fork.

Multi-Agent Operations Platform

Matrix-native operators · real infra tooling

Production multi-agent system: infrastructure ops, evaluation, project tracking, and workflow automation. Agents have real tools against live systems — not a demo chatbot.

LLM Evaluation Framework

GPQA-Diamond · IFEval · MATH · Custom agentic suite

Industry benchmarks (lm-evaluation-harness) plus custom agentic tests: tool-calling, schema/format, multi-turn, text quality. Independent LLM-as-judge. Dual-axis scorecard: quality and decode, not vanity tok/s. Grafana dashboards.

Serving & LLMOps

llm-switch · vLLM · SGLang · canary vs daily ship

Inference routed through llm-switch across vLLM and SGLang. Canaries on the same cluster as production serving. Speculative decode / MTP and quant recipes measured against a ship bar: quality up, or quality and speed. Documented rollback to the daily ship.

Persistent Agent Memory

Document banks · retain/recall · coverage

Long-lived semantic memory across agent sessions, with retain/recall and downtime backfill — not a one-shot RAG demo.

Model Fine-Tuning Pipeline

ORPO · LoRA · MoE

Preference alignment and parameter-efficient fine-tuning on hybrid architectures. Pre/post eval for catastrophic forgetting. Internal guidelines for small-active-parameter MoE.

ACTIVE STACK
Linux Proxmox ZFS Docker Kubernetes Terraform Python FastAPI pgvector Caddy nginx WireGuard Tailscale Unifi VPS Prometheus Grafana llm-switch vLLM SGLang NVFP4 1M-context VLM ORPO/LoRA deepseek-v4-flash deepseek-v4-flash-vision glm-5.3-flash Qwen3.5-397B

Open Source


Agentic-Backend

Production-pattern RAG stack — FastAPI + pgvector + Redis + Celery

Python

vllm-custom-main

vLLM fork with DGX Spark inference patches

Python

Agentic-HomeLab

Agentic automation for infrastructure management

Python

critical_user_journeys

2

CUJ framework reference for production SRE

Markdown
CERTIFICATIONS AWS DevOps ProfessionalCCNAGenAI Capstone — Agent FrameworksCompTIA Network+Dell DCSECisco DCUCI
Speaker — DevOps 207, Portland ME · Computer Network Information Systems, Wentworth Institute of Technology