Synapolis/Agent Substrate API Catalog
Дата проверки: 2026-08-07 UTC. Статус: рабочий каталог для инженерного выбора API substrate для "нагих агентов". Цены, лимиты, модели и политики быстро меняются; перед закупкой или подключением production-контуров нужно перепроверять live pricing/docs provider-а.
Назначение[edit | edit source]
Эта страница сравнивает LLM/chat/completion, code, tool-use/agents, embedding/rerank/search, browser/web, speech/vision, OpenRouter/BYOK/local/self-hosted и близкие substrate-слои, которые можно подключать к Synapolis agents.
Граница безопасности: каталог не является инструкцией по обходу закона, модерации или технических защит. "Low-policy / uncensored" ниже означает только сравнительную мягкость product policy, open-weight/self-hosted контроль и связанные юридические, операционные и security-риски.
Быстрый выбор для Synapolis[edit | edit source]
| Контур | Рекомендуемый substrate | Почему | Избегать |
|---|---|---|---|
| Дешёвые массовые фоновые агенты | Gemini Flash/Flash-Lite через Gemini API/Vertex; DeepSeek Flash/Pro для low-cost reasoning; Qwen/Together/Fireworks/Groq open-weight маршруты; локальный Ollama/vLLM для простых приватных задач. | Низкая цена за 1M токенов, достаточное качество для классификации, резюме, простых tool calls, фонового мониторинга. | Длинные автономные действия с секретами через BYOK-агрегаторы без ZDR; нестабильные preview-модели без fallback. |
| Coding agents | Anthropic Claude Sonnet/Opus для сложного репо и tool-use; OpenAI GPT-5-class/о-series/Codex surfaces; xAI Grok Build/Grok 4.5; Kimi/Qwen/DeepSeek/Together/Fireworks как price-performance альтернативы; локальный vLLM для private repo grep/patch с малой моделью. | Code tasks требуют контекста, tool-use дисциплины, стабильного JSON/function calling и низкой ошибки при правках. | Отправка приватных ключей, seed phrases, production .env и customer secrets в любые third-party prompts; auto-run shell без sandbox/gate. |
| Writing / analysis | OpenAI, Anthropic, Gemini Pro/Flash, Mistral, xAI, DeepSeek/Qwen/Kimi по цене и политике. | Нужны качество рассуждения, long context, стиль и устойчивость к галлюцинациям. | Search-less ответы по time-sensitive вопросам; модели без источников для юридических/финансовых/медицинских выводов. |
| Low-policy / offshore risk | Self-hosted open weights через vLLM/Ollama; отдельные Chinese providers/open models для price-performance, только на несекретных данных; OpenRouter как discovery, не как доверенный privacy boundary. | Меньше product-policy friction и больше control, но выше риск юрисдикции, supply-chain, censorship, logging ambiguity, export/sanctions/compliance. | Любой вред, незаконные цели, персональные данные без правового основания, secret-bearing workflows. |
| High-security / secret-sensitive | Локальный/self-hosted vLLM/Ollama/llama.cpp; cloud API только с ZDR/enterprise agreement, региональным processing, redaction, DLP, audit, separate keys, no prompt logging. | Секреты должны не покидать доверенный контур; browser/search substrate должен быть изолирован от credentials. | Public SaaS chat, free tiers, BYOK aggregators, browser agents с доступом к internal pages и внешнему web одновременно. |
| Search/browser/RAG substrate | Brave Search API для independent index/цитат; Tavily для agent search/extract/research; SerpAPI/SearchAPI для Google SERP fidelity; Perplexity Sonar для answer-with-citations; Voyage/Jina/Cohere/OpenAI embeddings+rerank для RAG. | Search substrate снижает hallucination, но вводит prompt injection и copyright/ToS риски от найденных страниц. | Автоматическое выполнение инструкций из web pages; ingestion без robots/ToS/license фильтра; хранение copyrighted corpora без policy. |
Сравнительная таблица[edit | edit source]
Сокращения: ZDR — zero data retention; BYOK — bring your own key; MTok — 1 млн токенов; официально означает, что строка опирается на публичные docs/pricing/policy provider-а, но не заменяет contract review.
| API/provider | Типы возможностей | Сильные стороны | Слабые стороны | Цена/стоимость ориентир и единица | Качество / latency / context / tool-use / code | Policy / copyright / content limits | Data retention / training / privacy controls | Secret leakage risk | Lock-in / availability / geography | Recommended use in Synapolis agents | Not recommended for | Integration notes | Source links / date checked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenAI | LLM/chat/responses, reasoning, tool-use, code, embeddings, speech/vision/image, batch, regional processing. | Сильная универсальность, mature API/SDK, хорошее tool/function calling, embeddings дешёвые, enterprise privacy/ZDR options. | Premium frontier costs; policy enforcement; модельный роутинг и названия быстро меняются; некоторые endpoints/features могут иметь особые retention limits. | От ~$0.05/MTok input на nano/cheap tiers до premium frontier; embeddings text-embedding-3-small около $0.02/MTok, 3-large около $0.13/MTok; batch обычно дешевле. Проверять live pricing. | Высокое качество для agents/code/reasoning; latency зависит от model/service tier; context у новых семейств large. | Строгие universal usage policies; ownership terms assign output to user where law allows; copyright/IP risk всё равно остаётся у интегратора по downstream use. | API data не используется для training без opt-in; default retention до 30 дней для abuse/service, ZDR для qualifying org/endpoints; regional processing с uplift для eligible models. | Средний: прямой API лучше BYOK-агрегатора, но secrets всё равно нельзя слать; включать redaction и project-level keys. | US provider; Azure/OpenAI/AWS Bedrock/regional options уменьшают, но не убирают lock-in. | Default high-quality substrate; code review, synthesis, structured tool-use, embeddings baseline. | Неподготовленные secret-bearing prompts; use cases без human oversight в regulated/high-stakes областях. | Использовать Responses API/tool schemas, batch for background, cache, service tiers; отдельные org/project keys. | pricing, data controls, usage policies, terms — checked 2026-08-07. |
| Anthropic Claude | LLM/chat/messages, extended thinking, tool-use, computer/code workflows, managed agents/Claude Code surfaces. | Очень сильный coding/repo reasoning, long-form analysis, cautious tool-use, enterprise posture. | Дороже low-cost конкурентов; policy stricter; model availability/rate limits can bottleneck; consumer Claude Code privacy differs from API/enterprise. | Claude pricing page: Sonnet 5 introductory ~$2/$10 per MTok input/output through 2026-08-31 then ~$3/$15; Haiku lower, Opus/Fable higher. Проверять live. | High quality for code agents and analysis; context/model family varies; latency often higher than small models. | Strong AUP/safety; явные disallowed/high-risk use constraints; copyright litigation/news risk exists, but API output/IP terms need contract review. | API commercial products: not used for training by default; standard deletion within 30 days except files/violations/law/agreements; ZDR available for eligible API/enterprise. | Средний-низкий при API+ZDR; высокий для consumer/free/chat or prompts containing repo secrets. | US provider, strong AWS/GCP/Bedrock routes; potential lock-in through Claude Code workflows. | Primary substrate for serious coding agents, careful writing/analysis, high-risk tool plans with review. | Cheap background classification; low-policy use; secret-bearing workflows without enterprise/ZDR. | Prefer direct API/Bedrock/Vertex where contracts fit; use prompt caching/batch; tool boundaries explicit. | pricing, models, retention, training, usage policy — checked 2026-08-07. |
| Google Gemini / Vertex AI / Agent Platform | Gemini API, Vertex AI, multimodal text/image/video/audio, grounding with Google Search, agents, embeddings, speech/vision via Google Cloud. | Strong multimodal, large context, cheap Flash tiers, Google Cloud IAM/VPC/region/data tooling, integrated Search grounding. | Policy/terms differ between free AI Studio, paid Gemini API, Vertex/Agent Platform; store=true defaults and logging settings need care. | Gemini API snippets show Flash-like paid tiers around $0.15/$1.25 MTok and higher Pro tiers; Vertex/Agent Platform Gemini 3.x Pro/Flash pricing can be $1.5-$4 input and $9-$18 output depending model/context/region. Search grounding after free quota about $14/1K requests. | Very good latency/cost on Flash; Pro stronger for analysis; excellent multimodal; tool-use acceptable but integration choices matter. | Google generative AI prohibited use policy; stricter around illegal/harmful, privacy, deception, explicit sexual content depending product. | Paid Gemini API: no training on prompts/responses per ZDR docs; ZDR requires avoiding store/session/logging features; Agent Platform request-response logging disabled by default but Interactions API may default store=true unless set false. | Средний: Google Cloud controls can be strong, but accidental storage/logging is easy. | Strong lock-in if using Vertex/Agent Platform/Search grounding; good regional/cloud availability. | Cheap massive background agents, multimodal ingestion, grounded search, GCP-native secure deployments. | Low-policy needs; workflows that cannot tolerate Google account/project policy changes. | Explicitly set store=false where applicable; separate AI Studio/free from paid API; use IAM/service accounts. | Gemini API pricing, Agent Platform pricing, Gemini ZDR, Agent Platform ZDR, policy — checked 2026-08-07. |
| xAI Grok | LLM/chat/code, reasoning, tool-use, image/video/voice APIs. | Grok 4.5 marketed for code, non-hallucination and agentic tool calling; comparatively lenient brand positioning; OpenAI-compatible style. | Smaller enterprise/privacy track record than OpenAI/Google/Anthropic; policy and terms changed recently; limited independent operational history. | Official pricing: grok-4.5 short context about $2 input / $6 output MTok, long context $4/$12; grok-build lower; media priced per image/sec. | Potentially strong code/agentic tasks; 500K+ context on flagship; latency depends on capacity. | AUP says maximize user control while requiring legal/responsible/safe use; not "uncensored" for illegal/harmful use. | Privacy policy applies; enterprise terms for API/business; verify API retention/training in contract/docs before secrets. | Средний-высокий until org-specific privacy terms are confirmed. | US provider; model catalog narrower; availability/capacity may fluctuate. | Experimental coding agents, writing/analysis alternative, low-friction public-data tasks. | Secret-sensitive or regulated workflows without enterprise terms; sole provider dependency. | Use as routed option behind gateway with fallback; log model/version and policy date. | pricing, models, AUP, enterprise terms, privacy — checked 2026-08-07. |
| Mistral AI | LLM/chat, agents Studio/Vibe, code Vibe, embeddings, OCR, fine-tuning, EU enterprise. | EU provider, open/proprietary mix, good price-performance, privacy controls, self-hostable/open-weight ecosystem. | Public pricing page may require docs/console for detailed model rows; ZDR only on Scale/stateless API; some products stateful. | API starts low; exact model prices vary. Official docs/pricing should be checked per model; third-party summaries place cheap tiers around $0.10+/MTok and frontier higher. | Good multilingual/code/general; latency good on smaller models; tool-use maturing. | Mistral usage/legal docs; generally less restrictive than Anthropic/OpenAI but not law-free. | API privacy controls for training/retention; ZDR available only Scale and stateless calls, not Vibe/agents/conversations/batch/stateful products. | Средний-низкий for Scale ZDR stateless; средний for normal API; high for stateful agent products with secrets. | EU geography advantage; lock-in moderate if using Studio/Vibe agents. | EU-friendly general agents, cheaper writing, open-weight/local transition path. | High-security unless ZDR/contract in place; tasks requiring best-in-class coding certainty. | Prefer stateless API for private use; document if using Labs/free tiers. | pricing, models, privacy controls, ZDR, terms — checked 2026-08-07. |
| Cohere | Command LLMs, enterprise RAG, Embed, Rerank, Model Vault/dedicated deployments. | Strong RAG/rerank/multilingual enterprise story; dedicated Model Vault for secure predictable inference. | Not usually first pick for frontier coding; pricing page mixes API and dedicated infra; generation catalog less buzz than frontier labs. | Rerank 4 Pro via OpenRouter listed ~$0.0025/search; dedicated Model Vault examples $4-$10/hour instances; older Embed/Rerank per-token varies by platform. | Excellent retrieval/reranking; Command useful for enterprise RAG; code weaker than Claude/OpenAI top tiers. | Enterprise acceptable-use terms; not a low-policy provider. | Enterprise privacy/security controls; verify per API/platform and cloud marketplace route. | Low-medium for dedicated/vault; medium for API; never send raw secrets into RAG rerank unless approved. | Canada/enterprise/cloud marketplace availability; some lock-in in RAG pipeline tuning. | RAG rerank, multilingual embeddings, private enterprise retrieval substrate. | Primary code agent model; cheap creative chat. | Use Cohere Rerank after cheap vector retrieval; cap docs/query and log token/search cost. | pricing, pricing docs, models, rerank price reference — checked 2026-08-07. |
| DeepSeek | LLM/chat/reasoning/code, OpenAI-compatible endpoint, long context, low-cost models. | Extremely strong price-performance for coding/reasoning; cheap background agents; large context on newer models. | Chinese jurisdiction/censorship/security concerns; pricing page warns of future increases; availability/payment/region risk. | Official docs: V4 Flash about $0.14 input cache miss / $0.28 output MTok; V4 Pro about $0.435/$0.87 MTok with very cheap cache-hit input; price increase warning. | Good code/reasoning for cost; latency/capacity can vary; tool-use depends wrapper. | Content/policy/censorship profile differs from US/EU providers; avoid political/legal-sensitive assumptions. | Privacy policy exists; do not assume Western enterprise ZDR. Treat as non-secret public-data substrate unless contract says otherwise. | Высокий for secrets/IP-sensitive prompts; acceptable for public/low-risk tasks after redaction. | China-linked; sanctions/export/procurement/geography risk; strong lock-in if prompts tuned to DeepSeek quirks. | Cheap background, code draft, long-context public analysis, fallback in router. | Secrets, customer data, regulated workloads, geopolitical/policy-sensitive analysis without cross-check. | Use via LiteLLM/OpenAI-compatible route with strict redaction; verify model IDs and peak pricing. | pricing, privacy — checked 2026-08-07. |
| Qwen / Alibaba Cloud Model Studio | Qwen text/code/multimodal/audio/video via Alibaba Cloud Model Studio; open-weight Qwen locally or via providers. | Strong open-weight ecosystem, cheap and capable multilingual/code models, cloud + self-host path. | International availability/region/payment and policy/censorship concerns; docs/pricing by region; China-linked procurement risk. | Official Model Studio pay-as-you-go by model; third-party summaries cite Qwen Flash around $0.10/$0.40 MTok, Plus around $0.40/$2.40, 397B higher. Verify live Alibaba Cloud region. | Very good price-performance; Qwen Coder strong; latency depends host. | Alibaba Cloud policies; Chinese jurisdiction/censorship risk; not for illegal/harmful use. | Verify Alibaba Cloud data handling per account/region; open weights self-host avoid API retention but shift ops risk to Synapolis. | High for direct cloud with secrets unless contract; low if self-hosted and isolated. | China/Singapore/international region constraints; open-weight reduces lock-in. | Cheap code/background via open weights; self-host candidates for private tasks. | Sensitive Western customer data/direct secrets on Alibaba-hosted API. | Prefer self-host/open-weight when privacy matters; use cloud for public/low-risk throughput. | models, pricing, Model Studio — checked 2026-08-07. |
| Moonshot / Kimi | Kimi LLM/code/long-context, Kimi K-series; API and open-weight/aggregator routes. | Strong code/agentic benchmarks and long-context claims; open-weight Kimi variants broaden hosting options. | Direct official pricing/resources can be less enterprise-transparent; China-linked risk; top models may be expensive. | Kimi resources page: K3 about $3 input cache miss / $0.30 cache hit / $15 output MTok; K2.6 about $0.95 miss / $0.16 hit / $4 output; OpenRouter K3 about $2.50/$14. | Strong coding and long-horizon agentic workflows if model claims hold; latency depends host/provider. | Treat as low-policy only in the sense of less Western SaaS friction; still legal/policy and censorship risks. | Do not assume ZDR on direct API/aggregators; self-hosted open weights give better control. | High for secrets unless self-hosted and audited. | China-linked and aggregator-dependent availability; open weights reduce vendor lock-in. | Coding experiments, public repo agents, cost/performance benchmark lane. | Secret-sensitive, regulated, political/censorship-critical outputs without independent verification. | Prefer OpenRouter/Together/Fireworks only for non-secret eval; pin provider route for reproducibility. | Moonshot, Kimi K3 pricing, Kimi K2.6 pricing, OpenRouter K3 — checked 2026-08-07. |
| Together AI | Serverless inference, dedicated endpoints, GPU clusters, fine-tuning, embeddings/rerank/moderation, open models. | Broad open-model catalog, serverless-to-dedicated path, privacy settings/ZDR, good for routing experiments. | Cost varies by model; passthrough third-party model privacy can differ; GPU/dedicated economics require capacity planning. | Serverless per MTok by model; public summaries show ~$0.03-$7+/MTok; GPU clusters/dedicated by GPU-hour; check live pricing. | Good latency/throughput; quality follows selected model; supports many code/open models. | Provider policy plus third-party/open-model licenses; not a single uniform content policy. | Privacy settings can disable storage/training and enable ZDR; docs note third-party model providers/passthrough need attention. | Medium; lower with ZDR and direct provider route, higher with passthrough/unknown models. | US provider; open catalog reduces model lock-in, platform lock-in remains. | Open-model benchmark lane, cheap background, dedicated endpoints for stable workloads. | Secrets unless ZDR/contract and no passthrough ambiguity. | Use model allowlists, privacy settings, spend caps; record exact model/provider. | pricing, privacy/security, terms/ZDR, privacy — checked 2026-08-07. |
| Fireworks AI | Serverless/on-demand/reserved inference, open models, embeddings, fine-tuning, OpenAI/Anthropic-compatible APIs. | High-throughput open-model serving, ZDR posture for open models, good enterprise options, batch 50% pricing. | Exact prices per model/tier; prepaid billing; hosted model quality follows open model. | Official: serverless per MTok; embeddings as low as $0.008-$0.1/MTok by model/size; batch 50% of serverless; H100/on-demand pricing separate. | Good latency/throughput for open models; code quality depends Qwen/Kimi/DeepSeek/Llama selected. | Open-model licenses and Fireworks terms; not an illegal/uncensored bypass. | Privacy policy says no AI training without opt-in; docs describe ZDR for most services, Response API has retention caveat. | Medium-low for open-model ZDR; medium if response/conversation state is used. | US provider; platform lock-in moderate; self-host transition possible by using open weights. | Cheap public-data agents, code model hosting, batch inference, fine-tuned open models. | Highest security unless contract; workflows needing proprietary frontier quality. | Use standard/priority/fast tier explicitly; batch background jobs; avoid stateful retention for secrets. | pricing, serverless pricing, ZDR/data handling, privacy — checked 2026-08-07. |
| GroqCloud | Fast inference for selected open models, Whisper/STT, TTS, tool-use if model supports. | Very low latency/tokens-per-second, cheap small/medium open models, self-serve ZDR controls. | Narrower model catalog; not always best quality; rate limits/capacity can shape production behavior. | Official docs: Llama 3.1 8B ~$0.05/$0.08 MTok; Llama 3.3 70B ~$0.59/$0.79; GPT-OSS 120B ~$0.15/$0.60; Whisper by audio hour. | Excellent latency; quality follows model; context often 131K on listed models. | Acceptable-use/responsible AI policies; open-model licenses apply. | Docs: usage metadata collected; customer data not retained by default except opt-in/persistence/reliability; ZDR self-serve. | Medium-low with ZDR; still third-party cloud, no raw secrets. | US provider, specialized hardware; model choice/capacity lock-in. | Fast cheap background agents, streaming UX, STT transcription, low-latency classifiers. | Hard reasoning/code requiring frontier accuracy; secret workflows without isolation. | Use as latency tier behind router; fallback on quality-sensitive tasks. | pricing, models, data, legal — checked 2026-08-07. |
| Cerebras | Ultra-fast inference for selected open/frontier-ish models; developer and paid tiers. | Very high speed claims; useful for latency-sensitive agents and batch acceleration. | Smaller catalog; detailed per-model token rates may require console; enterprise maturity narrower than hyperscalers. | Official pricing page: free trial $5 credits; developer self-serve starts around $10; public docs emphasize flexible pricing. Historical public rates were ~$0.10/MTok 8B and ~$0.60/MTok 70B, verify live. | Very low latency/high throughput; quality follows model selection. | Standard AUP/contracts; open-model licenses. | Verify retention/training terms in contract/trust center before secrets. | Medium until ZDR/contract confirmed. | US provider, specialized infra; availability/catalog lock-in. | Latency-sensitive public-data agents, fast eval loops. | Secret/high-compliance use without terms; broad multimodal tasks. | Treat as speed lane in router; benchmark live latency with Synapolis prompts. | pricing, press/site, historical pricing context — checked 2026-08-07. |
| Perplexity Sonar | Search-grounded chat/completions, citations, answer engine, OpenAI-compatible API. | Useful search/browser substrate with citations and low integration cost. | More expensive than raw search; citation quality must be verified; not a browser automation replacement. | Official docs: Gateway billed per token at model rates; Sonar pay-as-you-go no subscription. Third-party summaries cite base ~$1/$1 MTok plus request/search fees; verify live model card. | Good for current factual answers; latency includes search; not best for code edits. | Must respect source copyright/ToS; not a license to copy pages. | Verify API data terms; search queries may reveal intent/secrets. | Medium-high: search prompts often contain sensitive business questions. | US provider; source index/model routing dependency. | Current research with citations, pre-RAG discovery, fact-check lane. | Secret/internal search queries; legal/medical/financial final answers without human/source review. | Always store cited URLs/date; use as input to human/agent verification, not final truth. | pricing, Sonar quickstart — checked 2026-08-07. |
| Tavily | Web search, extract, crawl/research endpoints for AI agents/RAG. | Agent-oriented search/extraction API; simple credit pricing; good for web research pipelines. | Credits can burn quickly on /research; source quality/copyright/prompt injection risks remain. | Free 1,000 credits/month; PAYG ~$0.008/credit; monthly ~$0.0075-$0.005/credit; basic/advanced/research consume different credits. | Latency/search quality good for agent use; not an LLM by itself. | Web content license/copyright responsibility remains with integrator. | Search queries/extracted pages sent to Tavily; verify privacy terms for sensitive topics. | Medium: queries can leak plans; never include credentials. | SaaS provider; index/extraction behavior lock-in moderate. | Search/RAG substrate, source discovery, monitoring. | Secret-bearing browsing; executing page instructions. | Sanitize queries; isolate browser/search tool from credentialed tools; cite sources. | pricing, credits, product — checked 2026-08-07. |
| Brave Search API | Independent web/news/images/LLM context/search API. | Own index, predictable pricing, citations/LLM context, privacy-friendly brand. | Not Google SERP; niche/technical coverage may differ; lower QPS on small plans. | Official: Search $5 per 1,000 requests, includes $5 monthly credits; enterprise custom. | Good raw search/grounding; latency and completeness vary by query. | Search result snippets/URLs only; downstream copyright/ToS still apply. | Queries go to Brave; privacy better than many search APIs but not secret-safe. | Medium-low for non-secret queries; medium for strategy queries. | Vendor/index lock-in; global web availability. | Default cheap independent search for agents and RAG. | Need exact Google SERP/ads/local fidelity; secrets in search terms. | Use LLM Context endpoint where token-efficient; keep citation URLs. | pricing/API, pricing explanation — checked 2026-08-07. |
| SerpAPI | Google/Bing/Yahoo/Maps/Shopping/etc SERP JSON. | High SERP fidelity, mature JSON schema, legal shield on paid tiers. | More expensive than Brave/SearchAPI; not AI-specific answer engine. | Starter $25/month for 1,000 searches; Developer $75/5,000; Production $150/15,000; Big Data $275/30,000. | Good SERP extraction; latency tied to live search. | Public search collection/legal shield does not license downstream content use. | Queries expose intent; account logs/retention need vendor review. | Medium. | US provider; Google SERP dependency. | SEO/Google-specific search tasks, monitoring exact SERP features. | Cheap high-volume generic search; secret queries. | Use only when exact SERP matters; otherwise Brave/Tavily. | pricing, product/legal shield — checked 2026-08-07. |
| SearchAPI.io | Google SERP and related search APIs. | Cheaper Google SERP alternative with SLA/legal protection on higher plans. | Smaller ecosystem than SerpAPI; plan semantics and enhanced speed need modeling. | Developer $40/month, about $4/1K; Production $100/month, about $3/1K; BigData $250/month, about $2.5/1K; Octo 1M $1.5/1K. | Good for SERP/RAG discovery; verify fields for required verticals. | Same SERP/copyright downstream caveats. | Queries can leak sensitive plans. | Medium. | Vendor + Google SERP dependency. | Cost-sensitive Google SERP substrate. | Non-Google broad web answer generation; secrets. | Benchmark against SerpAPI for needed SERP features before switching. | pricing, product — checked 2026-08-07. |
| Voyage AI | Embeddings, multimodal embeddings, rerankers, Batch API. | Excellent retrieval-specialist provider; cheap rerank token pricing and free initial rerank allowance. | Narrow scope; acquired/operated with MongoDB ecosystem; not a chat model. | Rerank-2.5 about $0.05/MTok, lite $0.02/MTok, first 200M rerank tokens free; older embeddings: voyage-3.5 $0.06/MTok, lite $0.02/MTok. | High RAG quality; latency good; context varies by model. | Retrieval content copyright remains integrator responsibility. | Check docs/contract; embeddings encode sensitive content, so treat vectors as derived sensitive data. | Medium: documents/queries sent to third party and vectors may leak semantics. | Vendor lock-in via vector dimensions/model behavior; manageable with abstraction. | RAG indexing/rerank for Synapolis docs, public corpora, non-secret knowledge. | Raw secrets, private keys, personal data without approval. | Store model/version/dimension; separate indexes by sensitivity; rotate if changing model. | pricing, models — checked 2026-08-07. |
| Jina AI | Embeddings, rerankers, Reader, DeepSearch, multimodal search foundation models. | Strong multilingual/multimodal retrieval, Reader for web-to-markdown/JSON, open research/models. | Pricing details may be dashboard/package-based; less enterprise standard than hyperscalers. | Reranker/Embedding APIs begin with 10M free tokens per new API key; packages beyond. Reader/token pricing varies; verify account pricing. | Good retrieval quality, especially multilingual/multimodal; not primary chat LLM. | Web reader/crawler must respect website ToS/copyright. | API docs need review for retention; embeddings/reader inputs are sensitive derived data. | Medium. | Germany/EU-facing but global SaaS; model/vector lock-in. | RAG, reader/extraction, multilingual retrieval. | Secrets in docs/pages; high-compliance without contract. | Use Reader in isolated pipeline; don't execute extracted instructions. | product, reranker, embeddings — checked 2026-08-07. |
| OpenRouter | Aggregated LLM gateway, 400+ models, routing/fallbacks, BYOK/paygo, ZDR filters. | Fast model discovery, unified API, fallback/routing, privacy controls can restrict ZDR providers. | Adds extra trust boundary; provider logging varies; model/provider route can change unless pinned; BYOK concentrates keys in gateway. | Pay-as-you-go per model, claims transparent pricing and many free models; check each model card. | Quality/latency depends provider route; routing can improve availability but reduce reproducibility. | Policies differ by underlying provider; OpenRouter is not a policy bypass. | Docs: prompt retention opt-in; provider logging differs; ZDR can be enforced globally/per model group/guardrail/request. | High if using BYOK with secrets or prompt logging; medium with ZDR/non-secret prompts. | Strong platform lock-in for routing/account credits; reduces model lock-in. | Non-secret evaluation, fallback router, public-data price benchmarking. | High-security/secret-sensitive direct workflows; regulated data unless contract and ZDR routing are audited. | Pin model+provider when reproducibility matters; enable ZDR; disable prompt logging; never route secrets. | models, pricing, quickstart, data, provider logging, ZDR — checked 2026-08-07. |
| LiteLLM Proxy | Self-hosted LLM gateway/router, virtual keys, spend tracking, budgets, provider abstraction. | Good Synapolis control plane candidate: central keys, budgets, model allowlists, audit, OpenAI-compatible interface. | Security-critical service; past/current vulnerabilities possible; misconfig can expose master keys or allow over-broad routing. | OSS/self-host infra cost; provider token costs pass through; spend tracking uses model cost map. | Latency overhead small but depends routing; quality depends model; supports many providers. | Underlying provider policies apply; LiteLLM itself does not make unsafe use safe. | Privacy depends deployment DB/logging/proxy config and underlying providers. | High if proxy breached; low-medium if self-hosted locked down with per-agent virtual keys and no prompt logging. | Reduces provider lock-in; creates internal gateway lock-in/config dependency. | Synapolis internal model gateway, cost control, per-agent budgets, provider failover. | Running exposed to public internet without auth, patching, DB hardening, secret management. | Put behind private network; rotate master key; virtual keys per agent; redact logs; update pricing map. | virtual keys, spend tracking, budgets — checked 2026-08-07. |
| Ollama | Local/cloud open-model runner, embeddings, tool-capable models, offline mode. | Simple local deployment, offline possible, no per-token API cost, data stays local if actually local. | Quality/latency constrained by local hardware; operational burden; cloud features differ from local privacy. | Software free; cost is hardware/electricity/ops. Cloud models/regions may have separate pricing. | Good for small/medium models and private preprocessing; frontier quality usually lower unless large hardware. | Open model licenses vary; self-hosting removes SaaS policy but not law/copyright obligations. | Local/offline: no third-party retention; cloud: check Ollama data handling. | Low if isolated local; medium if exposed API or cloud. | Minimal vendor lock-in; model format/library ecosystem lock-in low. | High-security local summarization/classification, embeddings, private repo pre-pass. | Frontier coding decisions alone; exposed unauthenticated local API. | Bind to localhost/private network; model allowlist; monitor VRAM/RAM; do not mix browser tools with secrets. | product/data, model library, GitHub — checked 2026-08-07. |
| vLLM | Self-hosted high-throughput LLM serving, OpenAI-compatible server, batching, prefix caching, quantization/distributed. | Production-grade serving engine for open weights; good throughput; keeps data in Synapolis-controlled infra. | Requires GPU ops, model security, patching, observability; not a model by itself. | OSS; cost is GPU/hour + storage + ops. RunPod/Lambda/CoreWeave/local GPUs determine spend. | Excellent throughput/latency when tuned; quality depends selected model. | Open model licenses and local policy enforcement are Synapolis responsibility. | No third-party API retention if fully self-hosted; logs and traces become local sensitive artifacts. | Low if isolated/hardened; high if internet-exposed or logs leak prompts. | Low vendor lock-in at model/API layer; hardware/cloud lock-in possible. | Secret-sensitive private inference, high-volume background tasks, self-hosted code/RAG models. | Tiny VPS without GPU; tasks needing proprietary frontier accuracy. | Use OpenAI-compatible endpoint behind internal gateway; disable prompt logs or protect them; patch promptly. | vLLM, GitHub, docs mirror — checked 2026-08-07. |
| Replicate | Hosted public/private ML models, image/video/audio/LLM predictions, webhooks. | Broad model marketplace, quick prototypes, official models with stable API. | Per-second compute can be costly; public model supply-chain and license quality varies. | Pay-as-you-go compute time; model pages show estimate; public models billed by run time/hardware or model-specific units. | Great for media/prototypes; LLM quality depends hosted model and cold starts. | Model licenses/copyright vary; user responsible for generated media/code use. | Check retention/privacy for predictions/files; not secret-safe by default. | Medium-high for uploaded media/private data. | Platform/model marketplace lock-in; export possible for open models. | Vision/media experiments, non-secret model prototypes. | Secret docs, production LLM substrate at scale without cost model. | Prefer official models for stable API; set webhooks carefully; avoid private keys in inputs. | pricing, billing, official models — checked 2026-08-07. |
| RunPod | GPU pods, serverless GPUs, clusters for self-hosted models/vLLM/Ollama. | Flexible GPU cost, serverless scale-to-zero, good for self-hosting and batch jobs. | Ops burden, cold starts, container security, spot/preemption risk, storage/network extra costs. | Official pricing per GPU/hour or serverless per second; serverless bills from worker start until fully stopped; pricing changed June 2026 with Flex/Active workers. | Quality depends served model; latency depends worker/cold start/GPU. | Self-hosted model licenses and local policy enforcement. | Data stays in chosen deployment/provider infra; still third-party GPU cloud, so encrypt and avoid raw secrets unless contract. | Medium; lower with encrypted volumes/private network; high on community/spot for secrets. | GPU provider lock-in moderate; Docker/vLLM reduces. | Self-hosted open models, experiments, batch inference, private-ish code/RAG when contract acceptable. | Highest-security secrets without dedicated secure cloud/contract. | Harden images, no secrets baked into containers, restrict network, clean volumes. | pricing, serverless pricing, pricing update — checked 2026-08-07. |
| AWS Bedrock / Azure AI Foundry / cloud marketplaces | Managed access to Anthropic, Meta, Mistral, Amazon, OpenAI and other models with IAM/private networking/regional controls. | Enterprise procurement, IAM, VPC/PrivateLink, audit, compliance, data residency options. | Higher complexity/cost; model availability by region; service-tier lock-in. | Bedrock on-demand per token, batch 50% lower for select FMs, provisioned/reserved tiers; exact per-model rates region-dependent. | Quality depends model; latency and quotas by region/tier. | Underlying model + cloud provider policies both apply. | Stronger enterprise controls possible; configure logs carefully. | Low-medium with private networking/ZDR/contract; medium if default logging broad. | Strong cloud lock-in; better geography controls. | High-security enterprise route when Synapolis standardizes on cloud IAM/contracts. | Casual cheap experiments if direct API cheaper and data non-sensitive. | Use when compliance/procurement matters more than cheapest token. | Bedrock pricing, model availability — checked 2026-08-07. |
Security notes for agent substrate[edit | edit source]
- Secrets in prompts are exfiltration, not "context". Treat API keys, seed phrases, bearer tokens, SSH keys, cookies, private repo credentials, customer data and unpublished business plans as out-of-scope for third-party API calls unless the contour explicitly allows it.
- BYOK aggregators reduce integration work but increase blast radius: the aggregator can see prompts and may hold provider keys, route metadata, spend history and error traces. Use direct provider APIs or self-hosted gateway for sensitive work.
- Browser/search tools are prompt-injection surfaces. A web page can instruct the model to reveal memory, call tools, change code, or ignore policy. Search/extract output must be treated as untrusted data, not as system instructions.
- RAG vectors are derived sensitive data. Embeddings can leak semantic content and should be partitioned by sensitivity, model, tenant and deletion policy.
- Low-policy/open-weight does not remove law, copyright, privacy, export-control or platform abuse obligations. It only shifts enforcement from provider to Synapolis.
- For high-security agents: prefer local/self-hosted inference, no prompt logging, encrypted volumes, network egress allowlists, per-agent virtual keys, spend limits, and separate browser/search sandbox from credentialed tools.
Practical routing policy[edit | edit source]
| Workload | Default model lane | Cheap lane | Secure lane | Search/RAG lane |
|---|---|---|---|---|
| Simple background classify/summarize | Gemini Flash / OpenAI mini / Mistral small | DeepSeek Flash, Qwen Flash, Groq 8B/70B | Ollama/vLLM small model | Brave/Tavily only if current facts needed |
| Code patch/review | Claude Sonnet/Opus or OpenAI strong code/reasoning | Qwen/Kimi/DeepSeek via Fireworks/Together/Groq after tests | Local vLLM code model for private repo pre-pass; human-gated external call if needed | Repo-local grep/tests first; web search only for public docs |
| Long analysis/writing | OpenAI/Anthropic/Gemini Pro/xAI | DeepSeek/Qwen/Mistral | Local model for confidential drafts, then redacted external polish | Perplexity/Brave/Tavily with citations |
| Current research | Perplexity Sonar + verification LLM | Brave/Tavily + cheap summarizer | Self-hosted reader over approved sources | Voyage/Jina/Cohere rerank for local corpus |
| Secret-sensitive operation | No third-party by default | No cheap external lane | Local/self-hosted or enterprise ZDR/contract only | Internal docs only; no open web mixed with credentials |
Sources checked[edit | edit source]
Primary official docs/pricing/policy pages were preferred. Non-official pages were used only where official pages are dynamic, model-card-specific, or sparse, and the table marks those as orientation rather than contract source.
- OpenAI: pricing, data controls, usage policies, terms.
- Anthropic: pricing, retention, training, usage policy.
- Google: Gemini API pricing, Agent Platform pricing, Gemini ZDR, Agent Platform ZDR, policy.
- xAI: pricing, models, AUP, privacy.
- Mistral: pricing, models, privacy, ZDR.
- Cohere: pricing, pricing docs, models.
- DeepSeek: pricing, privacy.
- Alibaba/Qwen: models, pricing, Model Studio.
- Moonshot/Kimi: Moonshot, K3 pricing, K2.6 pricing, OpenRouter model card.
- Together: pricing, privacy/security, terms, privacy.
- Fireworks: pricing, serverless, data handling, privacy.
- Groq: pricing, models, data, legal.
- Cerebras: pricing, site, historical context.
- Search/RAG: Perplexity pricing, Sonar, Tavily pricing, Tavily credits, Brave Search API, SerpAPI pricing, SearchAPI pricing, Voyage pricing, Jina reranker, Jina embeddings.
- Aggregators/local: OpenRouter pricing, data collection, provider logging, ZDR, LiteLLM virtual keys, spend tracking, Ollama, Ollama library, vLLM, vLLM GitHub, Replicate pricing, billing, RunPod pricing, serverless pricing, Bedrock pricing.