Mark Mercado
Senior Platform Engineer | Cloud Infrastructure | Linux, Automation, and AI Systems
GitHub | LinkedIn | Portfolio | Resume | CloudMason
Summary
Senior platform and systems engineer with 30+ years of experience building the infrastructure behind reliable software: Linux operations, cloud automation, CI/CD, observability, HPC, and platform tooling. Currently focused on platform engineering at DigitalOcean, with prior depth in hybrid cloud, bare-metal infrastructure, high-performance computing, storage, networking, and operational automation.
Strong fit for infrastructure teams building AI, GPU, trading, or developer platforms: comfortable moving between low-level Linux troubleshooting, Kubernetes and automation workflows, source- controlled operations, and the human-facing documentation that keeps complex systems operable.
Selected Impact
- Built and operated cloud infrastructure automation around Ansible AWX, StackStorm, GitHub Actions, Kubernetes, internal APIs, and source-controlled operational workflows.
- Modernized infrastructure code by moving bespoke automation toward supported upstream collections, staging-first validation, repeatable runbooks, and safer rollout patterns.
- Built test-driven infrastructure readiness workflows for data center and region configuration, using canonical source data and local validation before production changes.
- Created observability systems with Prometheus, Grafana, StatsD-style metrics, Zabbix, custom exporters, and dashboards used for troubleshooting and operational review.
- Operated high-performance and hybrid infrastructure across HPC clusters, InfiniBand networks, Linux fleets, storage systems, virtualization, and cloud environments.
- Built agent workflow knowledge systems around MCP, Obsidian, reusable prompts, session memory, and human-in-the-loop review for AI-assisted engineering.
Technical Strengths
Languages
Python | Go | TypeScript | JavaScript | Bash | SQL | Ruby | Perl
Platform & Cloud
Linux | Kubernetes | Docker | Terraform | Ansible | Chef | AWX | StackStorm | GitHub Actions | CI/CD | DigitalOcean | AWS | bare-metal infrastructure
Reliability & Operations
Prometheus | Grafana | StatsD | Zabbix | PostgreSQL | Patroni | HAProxy | Nginx | DNS | SSO/SAML | runbooks | incident response | capacity and readiness validation
Systems & HPC
HPC clusters | InfiniBand | Lustre | ZFS | NetApp | virtualization | storage | networking | performance troubleshooting | RHEL/Solaris/UNIX
Developer & Agent Workflows
MCP | AI-assisted development | multi-agent workflows | evaluation and review loops | Obsidian knowledge systems | Markdown-driven documentation
Product & Web
React | Next.js | Node.js | REST APIs | component architecture | static sites | technical content systems
Experience
DigitalOcean - Platform Engineering
Systems / Platform Engineer, 2020 - Present
- Build and improve internal automation workflows for cloud infrastructure operations, spanning orchestration, configuration management, CI/CD, internal APIs, and operational tooling.
- Develop structured approaches for safe infrastructure changes: staging validation, reusable runbooks, source-controlled configuration, testable readiness checks, and reviewable automation.
- Work with Linux and Kubernetes-oriented platform systems where reliability depends on clear operational contracts, observable behavior, and repeatable rollout paths.
- Contribute to durable engineering knowledge systems that capture operational context and make future changes easier to reason about.
Barracuda Networks - Cloud / Infrastructure Operations
Lead Site Reliability Engineer, 2018 - 2020
- Operated hybrid cloud infrastructure across on-premises data centers and public cloud environments, including Linux fleets, private cloud, storage, networking, and service platforms.
- Implemented and maintained systems for bare-metal provisioning, image pipelines, DNS automation, package delivery, observability, and service discovery.
- Built Terraform, Terragrunt, Packer, Puppet, Ansible, Jenkins, Kubernetes, Helm, and Docker workflows for infrastructure delivery.
- Created Prometheus exporters, Grafana dashboards, operational runbooks, and deployment patterns for internal platform teams.
TotalCAE - High Performance Computing
Senior HPC Consultant, 2016 - 2018
- Administered a 1,200+ node HPC environment supporting engineering simulation workloads.
- Worked with RHEL, Lustre, InfiniBand, PBS Pro, Ansible, Zabbix, Elastic Stack, InfluxDB, Grafana, and custom operational integrations.
- Supported cluster operations across compute, storage, scheduler, network, monitoring, and user workload troubleshooting.
- Built automation and dashboards for cluster operations, reporting, and SLA-oriented support.
University of Michigan-Flint
UNIX Systems Administrator / Lecturer / Business Systems Analyst, 2005 - 2017
- Administered UNIX/Linux, Solaris, virtualization, storage, Oracle-backed ERP systems, LMS platforms, SSO, disaster recovery, monitoring, and CI/CD services.
- Built custom integrations in PL/SQL, Perl, PHP, Bash, and Python for campus systems and operational workflows.
- Taught computer science courses including Python, C++, operating systems, and programming fundamentals.
Early Web & Systems Work
Founding Partner / Systems Administrator / Web Developer, 1996 - 2003
- Built web applications, publishing systems, billing tools, search systems, secure client portals, and small-business infrastructure.
- Worked across FreeBSD, Linux, Solaris, Windows, PHP, Perl, JavaScript, MySQL, networking, firewalls, and backups.
Selected Projects
CloudMason
- Personal infrastructure and engineering hub for home lab systems, GitOps-style network management, DNS, Cloudflare, UniFi, NextDNS, runbooks, and operational experiments.
AgentBrain / OpenPaw / envx
- Personal agent workflow ecosystem exploring MCP, Obsidian-backed memory, machine convergence, reusable context, tool gateways, and repeatable AI-assisted engineering loops.
Public Technical Writing
- Maintains technical notes on CI/CD, SSH, secrets, Kubernetes, PostgreSQL, Patroni, Prometheus exporters, UniFi, Let's Encrypt, and home lab operations.
Education
Master of Science, Computer Science
University of Michigan-Flint
Bachelor of Science, Computer Science and Bachelor of Mathematics, Computer Science
University of Michigan-Flint
Publication
Cybersecurity in Banking and Financial Sector: Security Analysis of a Mobile Banking Application 2013 International Conference on Collaboration Technologies and Systems
Working Style
- Automation first, with enough documentation for future operators to understand the system.
- Comfortable moving from low-level infrastructure detail to product-facing developer experience.
- Strong preference for source-controlled workflows, explicit review, safe rollout, and observability.
- Interested in roles involving platform engineering, cloud infrastructure, Linux systems, reliability, AI infrastructure, GPU/HPC platforms, and AI-assisted engineering systems.