{"id":182,"date":"2026-09-23T11:40:26","date_gmt":"2026-09-23T11:40:26","guid":{"rendered":"https:\/\/www.dobryakov.net\/blog\/182\/"},"modified":"2026-09-23T11:40:26","modified_gmt":"2026-09-23T11:40:26","slug":"ai-governance-finops-resilience","status":"publish","type":"post","link":"https:\/\/www.dobryakov.net\/blog\/182\/","title":{"rendered":"FinOps &#038; Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over"},"content":{"rendered":"<p>A B2B SaaS team turned on semantic caching with a similarity threshold of 0.88 and a global namespace to save on tokens. The cache returned a Tier-1 customer&#x27;s cancellation summary during a Tier-3 customer&#x27;s session. Two different requests were &quot;similar enough&quot; by embedding distance \u2014 and a cost optimization became a cross-tenant data leak. The team rolled back to exact-match caching.<\/p>\n<p><!--more--><\/p>\n<p>That story is why the cost and resilience plane cannot be built in isolation from access control and privacy. An optimization that saves money at the cost of a leak is not savings. As a platform engineer, you own two things here: the bill and the availability. Uncontrolled token spend growth, exhausted API rate limits, and provider outages are your problem \u2014 and the solution is a single architectural layer that handles all of them.<\/p>\n<h2>The business promises you have to keep<\/h2>\n<p>For a company like Kovcheg (the reference platform for this course), the stakes are concrete. A provider outage during business hours means idle operations staff. An uncontrolled agent means a six-figure overnight bill. An exhausted rate limit means &quot;the assistant is down&quot; at peak hour.<\/p>\n<p>The business expects three commitments from the platform:<\/p>\n<ol>\n<li><strong>Spend is predictable and capped<\/strong> by department and by user.<\/li>\n<li><strong>A provider outage or throttle does not bring the system down.<\/strong><\/li>\n<li><strong>Optimizations never cross access or quality boundaries.<\/strong><\/li>\n<\/ol>\n<p>If any of these three break, it is an incident you own.<\/p>\n<h2>Threats and drivers<\/h2>\n<p>Four forces are working against you:<\/p>\n<ul>\n<li><strong>OWASP LLM LLM10 (Unbounded Consumption):<\/strong> cost and denial-of-service via expensive requests. A malicious or runaway prompt can burn through a budget in minutes.<\/li>\n<li><strong>Vendor outage and lock-in:<\/strong> a single provider is a single point of failure. When they throttle or go down, you go down.<\/li>\n<li><strong>Rate limits (HTTP 429):<\/strong> a traffic spike exhausts your quota. The API returns 429, and your application has to decide what to do next.<\/li>\n<li><strong>Cross-tenant leaks via cache:<\/strong> the opening story. An optimization that becomes a privacy incident.<\/li>\n<\/ul>\n<h2>The architectural answer: a multi-provider AI gateway<\/h2>\n<p>All model calls go through one layer: a multi-provider AI gateway with semantic caching, budgets, fallback, and circuit breaking. This is the same gateway required in Chapter 1 for PII redaction. Bypassing it is a privacy hole. The planes \u2014 PII (Ch. 1), guardrails (Ch. 3), tracing (Ch. 4), and now FinOps and resilience \u2014 do not spawn separate proxies. They live in one.<\/p>\n<p>The engineering stack:<\/p>\n<table>\n<thead>\n<tr>\n<th>Layer<\/th>\n<th>Tools<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>AI Gateway<\/strong><\/td>\n<td>LiteLLM Proxy (with <code>redis-semantic<\/code> or <code>qdrant-semantic<\/code> cache), Portkey, Kong AI Gateway (semantic cache plugin), Cloudflare AI Gateway, TrueFoundry<\/td>\n<\/tr>\n<tr>\n<td><strong>Cache &amp; limits<\/strong><\/td>\n<td>Redis (semantic cache, rate limiting, budget counters); GPTCache as an alternative<\/td>\n<\/tr>\n<tr>\n<td><strong>Resilience<\/strong><\/td>\n<td>Envoy (circuit breaking), backoff-based retries<\/td>\n<\/tr>\n<tr>\n<td><strong>Observability<\/strong><\/td>\n<td>Prometheus + Grafana<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>Engineering implementation<\/h2>\n<h3>Step 1. A single gateway for all model calls<\/h3>\n<p>Every model call \u2014 whether from a chatbot, a document processor, or an autonomous agent \u2014 goes through the gateway. No exceptions, no direct API calls from application code. This is the choke point where you enforce budgets, apply caching, and trigger fallback. If a team can bypass it, none of the guarantees below hold.<\/p>\n<h3>Step 2. Semantic caching with tenant isolation<\/h3>\n<p>Embed the prompt, search by cosine similarity in Redis, and above the threshold return the cached answer without calling the LLM. The opening incident teaches the configuration lessons:<\/p>\n<pre><code class=\"language-yaml\">semantic_cache:\n  similarity_threshold: 0.97          # above the typical 0.88 \u2014 fewer false hits\n  namespace: &quot;{tenant_id}:{model}&quot;    # model-level + tenant isolation, never global\n  skip_if: [&quot;contains_pii&quot;, &quot;personalized&quot;]  # never cache private content\n  per_request_threshold_override: true<\/code><\/pre>\n<p>Three rules are non-negotiable. The similarity threshold stays high \u2014 0.97, not 0.88 \u2014 because false hits in a regulated response (rates, contract terms) are dangerous. The namespace is scoped by tenant and model, never global. And private or personalized content is skipped entirely.<\/p>\n<p>In production, semantic caching realistically covers 20\u201345% of traffic. That is a material saving \u2014 as long as it does not break isolation.<\/p>\n<h3>Step 3. Multi-tenant budgeting<\/h3>\n<p>Hard limits on tokens and cost, enforced per department and per user. The progression is soft-alert to hard-stop: a department approaching its budget gets a warning, and at the limit, requests are rejected. For autonomous agents (Chapter 9), an additional per-task limit on steps and cost prevents a stuck loop from burning through an entire department&#x27;s allocation.<\/p>\n<h3>Step 4. Fallback and circuit breaker<\/h3>\n<p>When latency rises or the primary provider returns 5xx or 429 errors, the gateway auto-switches to a backup provider or degrades to a local model. Retries use exponential backoff. A circuit breaker opens the circuit to a failing provider instead of hammering it with requests that will only pile up and make the outage worse.<\/p>\n<h3>Step 5. Cost observability<\/h3>\n<p>Cost per request, per tenant, and per feature flows into Grafana. A spending anomaly triggers an alert. This is where a nighttime agent burning through budget gets caught \u2014 not when the monthly invoice arrives, but within minutes of the spike starting.<\/p>\n<h2>Where it breaks<\/h2>\n<p>Every mechanism in this chapter has a failure mode that turns the optimization into the incident.<\/p>\n<p><strong>Semantic cache returns &quot;almost the same.&quot;<\/strong> Requests close in embedding space but semantically different produce a wrong answer. In a regulated response \u2014 a loan rate, contract terms, a medical summary \u2014 that is dangerous. Use a high threshold and a cautious scope. For critical paths, exact-match caching only.<\/p>\n<p><strong>Cache versus privacy.<\/strong> A cached answer with someone else&#x27;s data is a leak. The opening case \u2014 Tier-1 cancellation data served to a Tier-3 user \u2014 is the canonical example. Tenant-scoped namespaces and skipping personal content are mandatory, not optional.<\/p>\n<p><strong>Fallback changes behavior.<\/strong> A different provider means different quality and different output format. Guardrails (Chapter 3) and evals (Chapter 7) must cover every provider in the fallback chain, or fallback silently degrades quality without anyone noticing until a customer complains.<\/p>\n<p><strong>Blind circuit breaking.<\/strong> An aggressive threshold cuts off legitimate traffic during a normal latency spike. Calibrate on real latency profiles, not on assumptions.<\/p>\n<p><strong>Cache masks drift.<\/strong> If 40% of answers come from cache, degradation in fresh answers (Chapter 7) is noticed later. The cache keeps serving old, correct answers while the live model&#x27;s quality silently drops. By the time someone notices, the drift is old news.<\/p>\n<h2>Standards and mapping<\/h2>\n<ul>\n<li><strong>OWASP LLM:<\/strong> LLM10 (Unbounded Consumption).<\/li>\n<li><strong>ISO\/IEC 42001:<\/strong> resource management, availability, continuity.<\/li>\n<li><strong>NIST AI RMF:<\/strong> Manage \u2014 resilience.<\/li>\n<li><strong>EU AI Act:<\/strong> robustness and continuity for high-risk systems (Art. 15).<\/li>\n<\/ul>\n<h2>Lab and artifact<\/h2>\n<p>Deploy LiteLLM or Portkey in front of the Kovcheg platform. Configure a semantic cache with a tenant-scoped namespace and a 0.97 threshold, skip for PII, set per-department budgets, and configure fallback to a backup provider.<\/p>\n<p>Reproduce the opening incident: set a global namespace and a 0.88 threshold, send two semantically close but distinct requests, and observe the cross-tenant leak. Then fix it via config and confirm the leak is gone. Simulate 429s and 5xxs from the primary provider and verify that degradation and circuit breaking work as expected.<\/p>\n<p>The artifact is a gateway config, a Grafana cost and latency dashboard, budget policies, and a &quot;cache hit-rate vs. cross-tenant safety&quot; report.<\/p>\n<p>The howto for this lab is at <a href=\"https:\/\/www.dobryakov.com\/howto\/ai-governance-finops-resilience.html\">https:\/\/www.dobryakov.com\/howto\/ai-governance-finops-resilience.html<\/a>.<\/p>\n<h2>Maturity checklist<\/h2>\n<table>\n<thead>\n<tr>\n<th>Level<\/th>\n<th>Capabilities<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>L1<\/strong><\/td>\n<td>A single gateway, basic limits, exact-match caching.<\/td>\n<\/tr>\n<tr>\n<td><strong>L2<\/strong><\/td>\n<td>Semantic caching with tenant isolation and skipping for private content, per-tenant budgets, fallback between providers.<\/td>\n<\/tr>\n<tr>\n<td><strong>L3<\/strong><\/td>\n<td>Circuit breaking, cost-anomaly alerts, evals across all providers, per-request threshold overrides, SLA monitoring.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/portkey.ai\/blog\/semantic-caching-thresholds\/\">Semantic caching thresholds and why they matter (Portkey)<\/a><\/li>\n<li><a href=\"https:\/\/www.getmaxim.ai\/articles\/top-semantic-caching-solutions-for-ai-applications-in-2026\/\">Top semantic caching solutions 2026 (Maxim)<\/a><\/li>\n<li><a href=\"https:\/\/neuraltrust.ai\/blog\/llm-caching-strategies\">LLM caching strategies (NeuralTrust)<\/a><\/li>\n<li><a href=\"https:\/\/noqta.tn\/en\/blog\/llm-gateway-multi-model-routing-guide-2026\">LLM gateway guide 2026<\/a><\/li>\n<\/ul>\n<p>A six-figure overnight bill and a cross-tenant leak come from the same mistake: treating cost optimization as a feature you bolt on, instead of a property of the gateway you build on. Get the gateway right \u2014 budgets, isolation, fallback \u2014 and production AI is a managed expense. Get it wrong, and you are one traffic spike away from a page that ruins your night.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A gateway with budgets, semantic caching, and fallback is the only thing standing between your AI workload and a six-figure overnight bill or a cross-tenant data leak.<\/p>\n","protected":false},"author":0,"featured_media":181,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[147,176,148],"tags":[177,181,179,180,178],"class_list":["post-182","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-governance","category-finops","category-llmops","tag-ai-gateway","tag-finops","tag-rate-limits","tag-resilience","tag-semantic-caching"],"aioseo_notices":[],"aioseo_head":"\n\t\t<!-- All in One SEO 4.9.10 - aioseo.com -->\n\t<meta name=\"description\" content=\"Build a multi-provider AI gateway with semantic caching, tenant isolation, budgets, and circuit breaking. Run AI in production without bankrupting yourself.\" \/>\n\t<meta name=\"robots\" content=\"max-image-preview:large\" \/>\n\t<link rel=\"canonical\" href=\"https:\/\/www.dobryakov.net\/blog\/182\/\" \/>\n\t<meta name=\"generator\" content=\"All in One SEO (AIOSEO) 4.9.10\" \/>\n\t\t<meta property=\"og:locale\" content=\"en_US\" \/>\n\t\t<meta property=\"og:site_name\" content=\"Grigoriy Dobryakov - IT+AI Blog - Grigoriy Dobryakov&#039;s blog: management, development and testing\" \/>\n\t\t<meta property=\"og:type\" content=\"article\" \/>\n\t\t<meta property=\"og:title\" content=\"AI FinOps &amp; Resilience: Cost, Rate Limits, High Availability\" \/>\n\t\t<meta property=\"og:description\" content=\"Build a multi-provider AI gateway with semantic caching, tenant isolation, budgets, and circuit breaking. Run AI in production without bankrupting yourself.\" \/>\n\t\t<meta property=\"og:url\" content=\"https:\/\/www.dobryakov.net\/blog\/182\/\" \/>\n\t\t<meta property=\"article:published_time\" content=\"2026-09-23T11:40:26+00:00\" \/>\n\t\t<meta property=\"article:modified_time\" content=\"2026-09-23T11:40:26+00:00\" \/>\n\t\t<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n\t\t<meta name=\"twitter:title\" content=\"AI FinOps &amp; Resilience: Cost, Rate Limits, High Availability\" \/>\n\t\t<meta name=\"twitter:description\" content=\"Build a multi-provider AI gateway with semantic caching, tenant isolation, budgets, and circuit breaking. Run AI in production without bankrupting yourself.\" \/>\n\t\t<script type=\"application\/ld+json\" class=\"aioseo-schema\">\n\t\t\t{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"BlogPosting\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/182\\\/#blogposting\",\"name\":\"AI FinOps & Resilience: Cost, Rate Limits, High Availability\",\"headline\":\"FinOps &#038; Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over\",\"author\":{\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/author\\\/#author\"},\"publisher\":{\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/#organization\"},\"image\":{\"@type\":\"ImageObject\",\"url\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/ai-governance-finops-resilience.jpg\",\"width\":1200,\"height\":630,\"caption\":\"FinOps & Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over\"},\"datePublished\":\"2026-09-23T11:40:26+00:00\",\"dateModified\":\"2026-09-23T11:40:26+00:00\",\"inLanguage\":\"en-US\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/182\\\/#webpage\"},\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/182\\\/#webpage\"},\"articleSection\":\"AI Governance, FinOps, LLMOps, AI gateway, FinOps, rate limits, resilience, semantic caching\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/182\\\/#breadcrumblist\",\"itemListElement\":[{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog#listItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/category\\\/ai-governance\\\/#listItem\",\"name\":\"AI Governance\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/category\\\/ai-governance\\\/#listItem\",\"position\":2,\"name\":\"AI Governance\",\"item\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/category\\\/ai-governance\\\/\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/182\\\/#listItem\",\"name\":\"FinOps &#038; Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over\"},\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog#listItem\",\"name\":\"Home\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/182\\\/#listItem\",\"position\":3,\"name\":\"FinOps &#038; Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over\",\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/category\\\/ai-governance\\\/#listItem\",\"name\":\"AI Governance\"}}]},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/#organization\",\"name\":\"Grigoriy Dobryakov - IT+AI Blog\",\"description\":\"Grigoriy Dobryakov's blog: management, development and testing\",\"url\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/182\\\/#webpage\",\"url\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/182\\\/\",\"name\":\"AI FinOps & Resilience: Cost, Rate Limits, High Availability\",\"description\":\"Build a multi-provider AI gateway with semantic caching, tenant isolation, budgets, and circuit breaking. Run AI in production without bankrupting yourself.\",\"inLanguage\":\"en-US\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/#website\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/182\\\/#breadcrumblist\"},\"author\":{\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/author\\\/#author\"},\"creator\":{\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/author\\\/#author\"},\"image\":{\"@type\":\"ImageObject\",\"url\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/ai-governance-finops-resilience.jpg\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/182\\\/#mainImage\",\"width\":1200,\"height\":630,\"caption\":\"FinOps & Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/182\\\/#mainImage\"},\"datePublished\":\"2026-09-23T11:40:26+00:00\",\"dateModified\":\"2026-09-23T11:40:26+00:00\"},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/\",\"name\":\"Grigoriy Dobryakov - IT+AI Blog\",\"description\":\"Grigoriy Dobryakov's blog: management, development and testing\",\"inLanguage\":\"en-US\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.dobryakov.net\\\/blog\\\/#organization\"}}]}\n\t\t<\/script>\n\t\t<!-- All in One SEO -->\n\n","aioseo_head_json":{"title":"AI FinOps & Resilience: Cost, Rate Limits, High Availability","description":"Build a multi-provider AI gateway with semantic caching, tenant isolation, budgets, and circuit breaking. Run AI in production without bankrupting yourself.","canonical_url":"https:\/\/www.dobryakov.net\/blog\/182\/","robots":"max-image-preview:large","keywords":"","webmasterTools":{"miscellaneous":""},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"BlogPosting","@id":"https:\/\/www.dobryakov.net\/blog\/182\/#blogposting","name":"AI FinOps & Resilience: Cost, Rate Limits, High Availability","headline":"FinOps &#038; Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over","author":{"@id":"https:\/\/www.dobryakov.net\/blog\/author\/#author"},"publisher":{"@id":"https:\/\/www.dobryakov.net\/blog\/#organization"},"image":{"@type":"ImageObject","url":"https:\/\/www.dobryakov.net\/blog\/wp-content\/uploads\/2026\/09\/ai-governance-finops-resilience.jpg","width":1200,"height":630,"caption":"FinOps & Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over"},"datePublished":"2026-09-23T11:40:26+00:00","dateModified":"2026-09-23T11:40:26+00:00","inLanguage":"en-US","mainEntityOfPage":{"@id":"https:\/\/www.dobryakov.net\/blog\/182\/#webpage"},"isPartOf":{"@id":"https:\/\/www.dobryakov.net\/blog\/182\/#webpage"},"articleSection":"AI Governance, FinOps, LLMOps, AI gateway, FinOps, rate limits, resilience, semantic caching"},{"@type":"BreadcrumbList","@id":"https:\/\/www.dobryakov.net\/blog\/182\/#breadcrumblist","itemListElement":[{"@type":"ListItem","@id":"https:\/\/www.dobryakov.net\/blog#listItem","position":1,"name":"Home","item":"https:\/\/www.dobryakov.net\/blog","nextItem":{"@type":"ListItem","@id":"https:\/\/www.dobryakov.net\/blog\/category\/ai-governance\/#listItem","name":"AI Governance"}},{"@type":"ListItem","@id":"https:\/\/www.dobryakov.net\/blog\/category\/ai-governance\/#listItem","position":2,"name":"AI Governance","item":"https:\/\/www.dobryakov.net\/blog\/category\/ai-governance\/","nextItem":{"@type":"ListItem","@id":"https:\/\/www.dobryakov.net\/blog\/182\/#listItem","name":"FinOps &#038; Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over"},"previousItem":{"@type":"ListItem","@id":"https:\/\/www.dobryakov.net\/blog#listItem","name":"Home"}},{"@type":"ListItem","@id":"https:\/\/www.dobryakov.net\/blog\/182\/#listItem","position":3,"name":"FinOps &#038; Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over","previousItem":{"@type":"ListItem","@id":"https:\/\/www.dobryakov.net\/blog\/category\/ai-governance\/#listItem","name":"AI Governance"}}]},{"@type":"Organization","@id":"https:\/\/www.dobryakov.net\/blog\/#organization","name":"Grigoriy Dobryakov - IT+AI Blog","description":"Grigoriy Dobryakov's blog: management, development and testing","url":"https:\/\/www.dobryakov.net\/blog\/"},{"@type":"WebPage","@id":"https:\/\/www.dobryakov.net\/blog\/182\/#webpage","url":"https:\/\/www.dobryakov.net\/blog\/182\/","name":"AI FinOps & Resilience: Cost, Rate Limits, High Availability","description":"Build a multi-provider AI gateway with semantic caching, tenant isolation, budgets, and circuit breaking. Run AI in production without bankrupting yourself.","inLanguage":"en-US","isPartOf":{"@id":"https:\/\/www.dobryakov.net\/blog\/#website"},"breadcrumb":{"@id":"https:\/\/www.dobryakov.net\/blog\/182\/#breadcrumblist"},"author":{"@id":"https:\/\/www.dobryakov.net\/blog\/author\/#author"},"creator":{"@id":"https:\/\/www.dobryakov.net\/blog\/author\/#author"},"image":{"@type":"ImageObject","url":"https:\/\/www.dobryakov.net\/blog\/wp-content\/uploads\/2026\/09\/ai-governance-finops-resilience.jpg","@id":"https:\/\/www.dobryakov.net\/blog\/182\/#mainImage","width":1200,"height":630,"caption":"FinOps & Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over"},"primaryImageOfPage":{"@id":"https:\/\/www.dobryakov.net\/blog\/182\/#mainImage"},"datePublished":"2026-09-23T11:40:26+00:00","dateModified":"2026-09-23T11:40:26+00:00"},{"@type":"WebSite","@id":"https:\/\/www.dobryakov.net\/blog\/#website","url":"https:\/\/www.dobryakov.net\/blog\/","name":"Grigoriy Dobryakov - IT+AI Blog","description":"Grigoriy Dobryakov's blog: management, development and testing","inLanguage":"en-US","publisher":{"@id":"https:\/\/www.dobryakov.net\/blog\/#organization"}}]},"og:locale":"en_US","og:site_name":"Grigoriy Dobryakov - IT+AI Blog - Grigoriy Dobryakov's blog: management, development and testing","og:type":"article","og:title":"AI FinOps &amp; Resilience: Cost, Rate Limits, High Availability","og:description":"Build a multi-provider AI gateway with semantic caching, tenant isolation, budgets, and circuit breaking. Run AI in production without bankrupting yourself.","og:url":"https:\/\/www.dobryakov.net\/blog\/182\/","article:published_time":"2026-09-23T11:40:26+00:00","article:modified_time":"2026-09-23T11:40:26+00:00","twitter:card":"summary_large_image","twitter:title":"AI FinOps &amp; Resilience: Cost, Rate Limits, High Availability","twitter:description":"Build a multi-provider AI gateway with semantic caching, tenant isolation, budgets, and circuit breaking. Run AI in production without bankrupting yourself."},"aioseo_meta_data":{"post_id":"182","title":"AI FinOps & Resilience: Cost, Rate Limits, High Availability","description":"Build a multi-provider AI gateway with semantic caching, tenant isolation, budgets, and circuit breaking. Run AI in production without bankrupting yourself.","keywords":null,"keyphrases":null,"primary_term":null,"canonical_url":null,"og_title":"AI FinOps & Resilience: Cost, Rate Limits, High Availability","og_description":"Build a multi-provider AI gateway with semantic caching, tenant isolation, budgets, and circuit breaking. Run AI in production without bankrupting yourself.","og_object_type":"default","og_image_type":"default","og_image_custom_url":null,"og_image_custom_fields":null,"og_image_url":null,"og_image_width":null,"og_image_height":null,"og_video":null,"og_custom_url":null,"og_article_section":null,"og_article_tags":null,"twitter_use_og":false,"twitter_card":"default","twitter_image_type":"default","twitter_image_custom_url":null,"twitter_image_custom_fields":null,"twitter_image_url":null,"twitter_title":"AI FinOps & Resilience: Cost, Rate Limits, High Availability","twitter_description":"Build a multi-provider AI gateway with semantic caching, tenant isolation, budgets, and circuit breaking. Run AI in production without bankrupting yourself.","schema_type":"default","schema_type_options":null,"schema":{"blockGraphs":[],"customGraphs":[],"default":{"data":{"Article":[],"Course":[],"Dataset":[],"FAQPage":[],"Movie":[],"Person":[],"Product":[],"ProductReview":[],"Car":[],"Recipe":[],"Service":[],"SoftwareApplication":[],"WebPage":[]},"graphName":"","isEnabled":true},"graphs":[]},"pillar_content":false,"robots_default":true,"robots_noindex":false,"robots_noarchive":false,"robots_nosnippet":false,"robots_nofollow":false,"robots_noimageindex":false,"robots_noodp":false,"robots_notranslate":false,"robots_max_snippet":null,"robots_max_videopreview":null,"robots_max_imagepreview":"large","priority":null,"frequency":null,"local_seo":null,"limit_modified_date":false,"ai":null,"breadcrumb_settings":null,"seo_analyzer_scan_date":null,"created":"2026-09-23 11:40:47","updated":"2026-09-23 11:40:47"},"aioseo_breadcrumb":"<div class=\"aioseo-breadcrumbs\"><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/www.dobryakov.net\/blog\" title=\"Home\">Home<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/www.dobryakov.net\/blog\/category\/ai-governance\/\" title=\"AI Governance\">AI Governance<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\tFinOps &amp; Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over\n\t\t<\/span><\/div>","aioseo_breadcrumb_json":[{"label":"Home","link":"https:\/\/www.dobryakov.net\/blog"},{"label":"AI Governance","link":"https:\/\/www.dobryakov.net\/blog\/category\/ai-governance\/"},{"label":"FinOps &#038; Resilience: How to Run AI in Production Without Going Bankrupt or Falling Over","link":"https:\/\/www.dobryakov.net\/blog\/182\/"}],"_links":{"self":[{"href":"https:\/\/www.dobryakov.net\/blog\/wp-json\/wp\/v2\/posts\/182","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.dobryakov.net\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.dobryakov.net\/blog\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/www.dobryakov.net\/blog\/wp-json\/wp\/v2\/comments?post=182"}],"version-history":[{"count":0,"href":"https:\/\/www.dobryakov.net\/blog\/wp-json\/wp\/v2\/posts\/182\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.dobryakov.net\/blog\/wp-json\/wp\/v2\/media\/181"}],"wp:attachment":[{"href":"https:\/\/www.dobryakov.net\/blog\/wp-json\/wp\/v2\/media?parent=182"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.dobryakov.net\/blog\/wp-json\/wp\/v2\/categories?post=182"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.dobryakov.net\/blog\/wp-json\/wp\/v2\/tags?post=182"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}