Arnaldo Sepulveda
Applied AI systems for Support Operations and enterprise workflows.
I investigate operational problems, establish what the evidence can support, and build and evaluate interventions ranging from process and deterministic software to retrieval, automation, and Applied AI.
I bring 12+ years of enterprise Customer Support and Technical Escalations experience from Genesys, where I worked across contact-center systems, knowledge, routing, integrations, customer and interaction data, production incidents, migrations, and supportability. Since late 2024, I have combined that operational foundation with hands-on Applied AI engineering through Keystone Applied Intelligence.
Operational intelligence and Applied AI engineering
Support Operations Intelligence
Evidence before intervention.
Support Operations Intelligence is a current portfolio and research project investigating how real operational evidence can be used to establish baselines, distinguish competing explanations, diagnose workflow constraints, compare interventions, and determine when Applied AI is actually warranted.
Current work is validating real Calgary 311 source evidence before moving into operational analytics and later event-level workflow analysis. This is the project direction; no completed workflow diagnosis, AI intervention, or operational outcome is claimed today.
Keystone Applied Intelligence
Keystone Applied Intelligence is where I build and test reference implementations for conversational AI, RAG, authorization-aware retrieval, agent workflows, and evaluation. The work spans Python and FastAPI APIs, PostgreSQL full-text search, pgvector, hybrid retrieval, deterministic reranking, evidence thresholds, local model serving, OpenTelemetry tracing, and debugging.
These public projects are separately composed engineering instruments, not one demonstrated production runtime. They let me test concrete mechanisms and retain internal evaluation evidence without treating one implementation or passing run as universal validation. Runtime governance is a secondary specialization within that broader AI systems work.
- Python APIs with FastAPI
- PostgreSQL FTS + pgvector retrieval
- RAG with evidence thresholds
- Authorization-aware retrieval
- Conversational and agent workflows
- Evaluation and OpenTelemetry tracing
Secondary research: Governed Execution. Its bounded Track A reference implementation, Runtime Validity, studies controlled process-local authority change, revalidation behavior, and transition evidence. It does not represent external revocation or production authorization. runtime-validity on GitHub.
Positions earned by implementation
The knowledge may already exist; the hard part is making it usable
In enterprise support, the answer often already exists somewhere: a prior case, documentation, an engineering discussion, or an experienced person’s memory. The harder problem is retrieving the right knowledge for the right person and context, with citations and authorization boundaries, without forcing another engineer to reconstruct the same reasoning.
Evaluation should be able to embarrass the system, not flatter it
An evaluation that cannot surface failures provides weak evidence. The useful ones are built to surface failure (adversarial access probes, out-of-scope queries, fail-closed cases), and the failing runs get published next to the passing ones. That is where the real design feedback comes from.
Production AI is a systems problem
A useful model is only one component. Production behavior also depends on retrieval, APIs, state, integrations, latency, escalation, observability, failure handling, and the surrounding operational workflow.
Some controls belong outside the prompt
When a requirement must hold regardless of model wording, it often belongs in deterministic system logic: retrieval filters, database predicates, state transitions, refusal rules, or authorization checks.
Local and customer-controlled infrastructure exposes engineering tradeoffs
Local inference makes tradeoffs in latency, model capacity, failure modes, operational ownership, and data boundaries explicit. Those constraints can clarify the architecture without making cloud APIs inherently inferior.
Notes from building AI systems
I write about retrieval, evaluation, enterprise AI engineering, and runtime controls: lessons that emerge from implementation, measurement, and operational experience.
How repeated enterprise support investigations pushed me toward retrieval, AI-assisted workflows, and systems that make existing organizational knowledge useful when people need it.
Retained failures, bounded claims, and the engineering feedback that a polished demonstration cannot provide.
How contact-center operations, escalation, integration, and incident response shape my approach to AI systems.
Enterprise engineering is the foundation, not a previous chapter
I spent 12+ years at Genesys working within enterprise contact-center systems. In Business Applications, I specialized in Knowledge and Knowledge Center, AI and classification systems, Digital Services, Agent Workspace, customer and interaction data, routing, conversational systems, and enterprise integrations. This work covered implementations, go-lives, migrations, high-severity incidents, distributed troubleshooting, customer-facing technical investigation, and production recovery.
I also troubleshot WFM-integrated agent and supervisor workflows and operational statistics, tracing missing or incorrect data across application and data boundaries. I worked directly with customers, product managers, developers, and technical directors on product behavior, supportability, customer requirements, and deployment architecture, including clustered and high-volume customer and interaction-data deployments. I later led the Genesys Cloud CX UI Support Team. My AI engineering work since late 2024 builds directly on that experience.
I hold an MScE in Electrical Engineering from the University of New Brunswick, with a thesis applying machine learning to smart-grid load control.
What is public now
I try to make claims that map to something you can inspect. The internal eval baselines below are published with their methodology and lineage in the ledger. They are commit-bound project evidence, not independent validation.
Governed retrieval, keystone-core/retrieval-v1 (2026-04-11): P@1 0.75, MRR 0.79,
adversarial ACL 8/8 blocked, fail-closed 5/6, Alberta OHS safety corpus.
Governed agent evaluation run, keystone-core/agent-v1 (2026-05-20): 186 cases
across 12 categories and 558 executions, with 153 strict passes, 33 characterization
cases, and 0 strict failures at the evaluated keystone-gov commit. The failing precursor
run is published alongside it.
Work I want to do next
I am interested in work where I can combine enterprise Support Operations experience, operational analytics, workflow diagnosis, and hands-on Applied AI engineering. That includes conversational AI, retrieval, AI-assisted workflows, evaluation, APIs, and enterprise integrations, as well as upstream work to determine which intervention the evidence actually supports. I am especially interested in systems used in real customer and operational workflows where outcomes can be measured, investigated, and improved. If that is the kind of problem you are staffing, I would like to hear about it.