Skip to main content

Kach Prompt ki Konsyan de Provider

Superdav AI Agent v1.12.0 prezante kach prompt ki konsyan de provider, ki optimize depans API ak latans lè li mete prompt yo an kach atravè diferan provider LLM. Chak provider gen mekanis ak konfigirasyon kach diferan.

Apèsi

Kach prompt pèmèt ou:

  • Mete gwo prompt yo itilize souvan an kach
  • Diminye depans API lè w evite pwosesis ki repete san nesesite
  • Amelyore latans pou demann ki an kach
  • Jere sik lavi kach la klèman

Diferan provider aplike kach yon fason diferan:

  • Google Gemini: cachedContents API
  • Azure OpenAI: Kach prompt ak TTL
  • OpenRouter: Kach espesifik pou provider
  • Vertex Anthropic: Kach prompt ak kontwòl kach

Google Gemini: cachedContents API

Google Gemini bay jesyon kach klè atravè cachedContents API.

Konfigirasyon

$config = [
'provider' => 'google-gemini',
'model' => 'gemini-2.0-flash',
'caching' => [
'enabled' => true,
'ttl' => 3600, // 1 hour in seconds
'max_tokens' => 1000000, // Max tokens to cache
],
];

Kreye yon Prompt ki an Kach

use Superdav\AI\Providers\GoogleGemini;

$gemini = new GoogleGemini( $config );

$cached_content = $gemini->create_cached_content(
[
'system_prompt' => 'You are a helpful assistant...',
'context' => 'Large context document...',
'ttl' => 3600,
]
);

// Returns: ['cache_id' => 'abc123', 'expires_at' => timestamp]

Itilize yon Prompt ki an Kach

$response = $gemini->generate(
[
'cache_id' => 'abc123',
'prompt' => 'User question here',
]
);

Sik Lavi Kach

// List cached contents
$caches = $gemini->list_cached_contents();

// Get cache details
$cache = $gemini->get_cached_content( 'abc123' );

// Extend cache TTL
$gemini->update_cached_content(
'abc123',
['ttl' => 7200] // Extend to 2 hours
);

// Delete cache
$gemini->delete_cached_content( 'abc123' );

Pi Bon Pratik pou Gemini

  • Mete TTL ki apwopriye: Balanse ekonomi depans ak kach ki ka vin demode
  • Mete system prompts an kach: Reyitilize menm system prompt la atravè demann yo
  • Siveye itilizasyon kach: Swiv ki kach yo itilize plis
  • Netwaye kach ki ekspire: Efase kach ki pa itilize yo detanzantan

Azure OpenAI: Kach Prompt

Azure OpenAI sipòte kach prompt ak jesyon TTL otomatik.

Konfigirasyon

$config = [
'provider' => 'azure-openai',
'model' => 'gpt-4-turbo',
'api_version' => '2024-08-01-preview',
'caching' => [
'enabled' => true,
'cache_control' => 'max_age=3600',
],
];

Aktive Kach

use Superdav\AI\Providers\AzureOpenAI;

$azure = new AzureOpenAI( $config );

$response = $azure->generate(
[
'system_prompt' => 'You are a helpful assistant...',
'context' => 'Large context document...',
'prompt' => 'User question here',
'cache_control' => 'max_age=3600',
]
);

// Response includes cache usage:
// [
// 'content' => '...',
// 'cache_creation_input_tokens' => 1000,
// 'cache_read_input_tokens' => 500,
// ]

Header Kach

Azure OpenAI itilize header HTTP pou kontwòl kach:

Cache-Control: max_age=3600

Valè ki sipòte yo:

  • max_age=<seconds>: Mete an kach pou dire ki espesifye a
  • no_cache: Pa mete demann sa a an kach
  • no_store: Pa mete an kach epi pa reyitilize

Siveye Itilizasyon Kach

$response = $azure->generate( [...] );

$cache_tokens = $response['cache_creation_input_tokens'] ?? 0;
$cache_hits = $response['cache_read_input_tokens'] ?? 0;

echo "Cache creation: $cache_tokens tokens\n";
echo "Cache hits: $cache_hits tokens\n";

Pi Bon Pratik pou Azure OpenAI

  • Itilize prompt ki konsistan: Prompt ki idantik benefisye de kach
  • Mete TTL ki rezonab: Balanse depans ak frechè
  • Siveye metrik kach: Swiv kreyasyon kach kont hits
  • Gwoupe demann ki sanble: Gwoupe demann yo pou maksimize hits kach

OpenRouter: Kach Espesifik pou Provider

OpenRouter sipòte kach atravè provider ki anba yo (OpenAI, Anthropic, elatriye).

Konfigirasyon

$config = [
'provider' => 'openrouter',
'model' => 'openai/gpt-4-turbo',
'caching' => [
'enabled' => true,
'provider_cache' => 'openai', // Use OpenAI's caching
],
];

Itilize Kach OpenRouter

use Superdav\AI\Providers\OpenRouter;

$router = new OpenRouter( $config );

$response = $router->generate(
[
'system_prompt' => 'You are a helpful assistant...',
'context' => 'Large context document...',
'prompt' => 'User question here',
'cache_control' => 'max_age=3600',
]
);

Opsyon Espesifik pou Provider

Diferan provider gen mekanis kach diferan:

// OpenAI-compatible caching
$response = $router->generate(
[
'model' => 'openai/gpt-4-turbo',
'cache_control' => 'max_age=3600',
]
);

// Anthropic-compatible caching
$response = $router->generate(
[
'model' => 'anthropic/claude-3-opus',
'cache_control' => [
'type' => 'ephemeral',
'max_tokens' => 1000000,
],
]
);

Pi Bon Pratik pou OpenRouter

  • Konnen kach provider ou a: Chak provider gen mekanis diferan
  • Teste konpòtman kach: Verifye kach la mache ak provider ou chwazi a
  • Siveye depans yo: Swiv ekonomi ki soti nan kach
  • Itilize modèl ki konsistan: Chanje modèl kraze hits kach yo

Vertex Anthropic: Kach Prompt ak Kontwòl Kach

Vertex Anthropic (Google Cloud) sipòte kach prompt ak kontwòl kach klè.

Konfigirasyon

$config = [
'provider' => 'vertex-anthropic',
'model' => 'claude-3-opus',
'project_id' => 'your-gcp-project',
'region' => 'us-central1',
'caching' => [
'enabled' => true,
'cache_control' => [
'type' => 'ephemeral',
'max_tokens' => 1000000,
],
],
];

Sèvi ak Vertex Anthropic Caching

use Superdav\AI\Providers\VertexAnthropic;

$vertex = new VertexAnthropic( $config );

$response = $vertex->generate(
[
'system_prompt' => 'You are a helpful assistant...',
'context' => 'Large context document...',
'prompt' => 'User question here',
'cache_control' => [
'type' => 'ephemeral',
'max_tokens' => 1000000,
],
]
);

// Response includes cache metrics:
// [
// 'content' => '...',
// 'usage' => [
// 'input_tokens' => 1000,
// 'cache_creation_input_tokens' => 500,
// 'cache_read_input_tokens' => 300,
// ],
// ]

Kalite Kontwòl Cache

  • ephemeral: Cache pou dire demann nan (default)
  • persistent: Cache atravè plizyè demann (si li sipòte)

Siveyans Itilizasyon Cache

$response = $vertex->generate( [...] );

$usage = $response['usage'];
$cache_created = $usage['cache_creation_input_tokens'] ?? 0;
$cache_read = $usage['cache_read_input_tokens'] ?? 0;

echo "Cache created: $cache_created tokens\n";
echo "Cache read: $cache_read tokens\n";

Pi Bon Pratik pou Vertex Anthropic

  • Sèvi ak caching ephemeral: Bon pou caching nan yon sèl sesyon
  • Mete max_tokens jan sa apwopriye: Balanse gwosè cache ak pri
  • Siveye metrik cache yo: Swiv efikasite cache la
  • Teste ak chaj travay ou: Verifye caching bay avantaj pou ka itilizasyon ou

Estrateji Caching Ant Founisè

Konfigirasyon Inifye

$config = [
'caching' => [
'enabled' => true,
'default_ttl' => 3600,
'providers' => [
'google-gemini' => [
'ttl' => 3600,
'max_tokens' => 1000000,
],
'azure-openai' => [
'cache_control' => 'max_age=3600',
],
'vertex-anthropic' => [
'cache_control' => [
'type' => 'ephemeral',
'max_tokens' => 1000000,
],
],
],
],
];

Deteksyon Founisè

$provider = $config['provider'];

$cache_config = $config['caching']['providers'][ $provider ]
?? $config['caching'];

// Use provider-specific caching configuration

Estrateji Fallback

try {
// Try caching with primary provider
$response = $primary_provider->generate( $request );
} catch ( CacheException $e ) {
// Fall back to non-cached request
$response = $primary_provider->generate(
array_merge( $request, ['cache_control' => 'no_cache'] )
);
}

Optimizasyon Pri

Kalkile Ekonomi

$cache_created_tokens = $response['cache_creation_input_tokens'] ?? 0;
$cache_read_tokens = $response['cache_read_input_tokens'] ?? 0;
$regular_tokens = $response['input_tokens'] ?? 0;

// Typical pricing (varies by provider):
$cache_creation_cost = $cache_created_tokens * 0.00001; // 10x cheaper
$cache_read_cost = $cache_read_tokens * 0.000001; // 100x cheaper
$regular_cost = $regular_tokens * 0.00001;

$total_cost = $cache_creation_cost + $cache_read_cost + $regular_cost;
$savings = ($regular_tokens * 0.00001) - $total_cost;

echo "Estimated savings: \$$savings\n";

Konsèy Optimizasyon

  • Cache gwo system prompts: Pi gwo ekonomi pri
  • Reitilize kontèks: Cache dokiman kontèks yo itilize souvan
  • Gwoupe demann yo: Mete demann ki sanble yo ansanm pou maksimize cache hits
  • Siveye efikasite cache la: Swiv ekonomi reyèl yo
  • Ajiste TTL: Balanse pri ak fraîcheur

Depanaj

Cache pa itilize

  • Verifye caching aktive nan konfigirasyon an
  • Tcheke prompts yo idantik (caching mande korespondans egzak)
  • Verifye cache la pa ekspire
  • Tcheke limit cache espesifik pou founisè a

Kreyasyon cache ap echwe

  • Verifye gwosè cache la nan limit founisè a
  • Tcheke sentaks kontwòl cache la kòrèk
  • Asire founisè a sipòte caching pou modèl ou a
  • Revize dokimantasyon founisè a pou limit yo

Pri inatandi

  • Siveye kreyasyon cache kont cache read tokens
  • Verifye cache la vrèman ap itilize
  • Tcheke pou cache misses akòz varyasyon nan prompt yo
  • Konsidere ajiste TTL oswa estrateji cache

Konparezon Founisè

FonksyonaliteGeminiAzure OpenAIOpenRouterVertex Anthropic
Cache APIcachedContentsHTTP headersEspesifik pou founisèKontwòl cache
Kontwòl TTLEksplisitAtravè headersDepann de founisèEphemeral/persistent
Gwosè cache maksimòm1M tokensDepann de founisèDepann de founisè1M tokens
Rediksyon pri90%90%Depann de founisè90%
SiveyansDetayeAtravè metrikDepann de founisèAtravè itilizasyon

Pwochen Etap yo

  1. Chwazi founisè ou: Chwazi selon bezwen ou yo
  2. Konfigire caching: Mete caching espesifik pou founisè a an plas
  3. Teste caching: Verifye li mache ak prompts ou yo
  4. Siveye itilizasyon: Swiv cache hits ak ekonomi pri
  5. Optimize: Ajiste TTL ak estrateji cache selon rezilta yo