Upgrade to Pro
— share decks privately, control downloads, hide ads and more …
Speaker Deck
Sign up for free
Menu
Search
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Features
All features
Private URLs
Password Protection
Custom URLS
Scheduled publishing
Remove Branding
Restrict embedding
Deck Collections
Notes
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Explore
Featured decks
Featured speakers
Programming
Technology
Storyboards
Pricing
Search
Sign in
Sign up for free
【re:Invent 2024 アプデ】 Prompt Routing の紹介
Search
Champ
December 17, 2024
Technology
580
1
Share
Embed
Copy iframe code
Copy JS code
Copy link
Start on current slide
【re:Invent 2024 アプデ】 Prompt Routing の紹介
Champ
December 17, 2024
More Decks by Champ
See All by Champ
MCPサーバー、AWSのどこに置く?
champ
0
160
Kiro CLI 徹底解剖
champ
0
48
Amazon Bedrockの自動推論チェックを検証!
champ
0
33
Amazon BedrockでClaude 3.5 Sonnet v2のComputer useを試す
champ
0
150
【Bedrock×Athena】生成系AIでSlackデータの分析に挑戦
champ
0
250
Amazon Qの全体像を掴んでみよう!
champ
0
99
神アプデ?Amazon Comprehendで 生成系AIの毒性検出に挑戦!
champ
0
410
Bedrockで挑戦! 生成系AIで Slackコミュニケーションの活性化!
champ
0
490
Other Decks in Technology
See All in Technology
『GOエコノミー 』(相乗りサービス) におけるスペック駆動開発
mot_techtalk
1
120
セルフサービスのオブザーバビリティ基盤をOpenTelemetryで作る / Building a Self-Service Observability Platform with OpenTelemetry
ymotongpoo
3
500
Claude起点の仕様駆動開発
tanakaseiya
0
260
Swap and Memory Reclaim - Squeezing Out More RAM
ennael
PRO
1
1.5k
いちAWSエンジニアのAI活用を振り返る #devio2026 / devio osaka 2026 kawahara
masahirokawahara
1
300
【データ横丁主催】AI Agentがコンテキストを使って仕事をした後、何が残るのか― 組織の経験を次の判断に引き継ぐ「Agent Memory」
shisyu_gaku
2
300
行動するAIのためのオントロジー | DevRev — Encraft #26.pdf
dvrv_tknrszk
2
650
CI/CDではもう遅い - 人とAIが迂回しないDevSecOps Verify基盤の再設計 -
kintotechdev
1
560
Confitura 2026
logico_jp
0
120
2026-10-01_MagicPod_QAハーネスエンジニアリングとQA組織の未来像
ynisqa1988
0
110
2026-09-26 Platform Engineering Kaigi 2026 インフラとアプリの境界線と委譲の設計 / Drawing the Infra and App Line
masasuzu
0
540
AIに書かせて、プラットフォームで縛る ― EKSプラットフォームで実践した責任境界と権限設計
elmodev09
1
1.1k
Featured
See All Featured
Why Mistakes Are the Best Teachers: Turning Failure into a Pathway for Growth
auna
0
310
SEO Brein meetup: CTRL+C is not how to scale international SEO
lindahogenes
2
2.9k
Navigating the moral maze — ethical principles for Al-driven product design
skipperchong
2
590
YesSQL, Process and Tooling at Scale
rocio
174
15k
JAMstack: Web Apps at Ludicrous Speed - All Things Open 2022
reverentgeek
1
620
Creating an realtime collaboration tool: Agile Flush - .NET Oxford
marcduiker
35
2.6k
Claude Code どこまでも/ Claude Code Everywhere
nwiizo
68
58k
A Modern Web Designer's Workflow
chriscoyier
699
190k
The Mindset for Success: Future Career Progression
greggifford
PRO
0
520
What the history of the web can teach us about the future of AI
inesmontani
PRO
1
720
HTML-Aware ERB: The Path to Reactive Rendering @ RubyCon 2026, Rimini, Italy
marcoroth
5
750
Future Trends and Review - Lecture 12 - Web Technologies (1019888BNR)
signer
PRO
0
3.8k
Transcript
【re:Invent 2024 アプデ】 Prompt Routing の紹介
自己紹介
1. (Intelligent)Prompt Routing とは 新機能の概要 re:invent 2024 で発表された新機能(プレビュー) プロンプトの複雑さを自動判定し、最適なモデルへ自動 振り分け
2024/12/17 時点では以下のルーティングが可能 Claude Sonnet 3.5 と Claude 3 Haiku Llama 3.1 70B と Llama 3.1 8B なにが嬉しいのか? プロンプトを適切なモデルにルーティングすることでコ ストを下げることが可能
2. 仕組み 処理の流れ Prompt Routing は以下の流れで処理が行われます プロンプト受信 パフォーマンスの計算 ルーティングの実施 実行とフォールバック
それぞれについて解説していきます
2. 仕組み プロンプト受信 Prompt Routing がプロンプトを受け取る プロンプトの特徴を分析(長さ、複雑さ、要求タスクなど)
2. 仕組み パフォーマンスの計算 設定された各モデル(例:Sonnet と Haiku)でのパフォーマンスを計算 推論は実行せず、パフォーマンスの計算のみを実施 モデル間の品質差(quality_difference)を算出
2. 仕組み ルーティングの実施 quality_difference と 閾値(responseQualityDifference) を比較 quality_difference が 閾値未満の場合、軽量モデルを選択
閾値以上の場合、高性能モデルを選択 2024/12/17 時点ではデフォルト値は 0.0 になっている? ので、差が少しでもあれば Sonnet を選択
2. 仕組み ルーティングの実装 続き 簡略化したルーティングロジックのイメージ quality_difference = high_quality_model_score - lightweight_model_score
responseQualityDifference = 0.1 # 閾値が 0.1 の場合 if quality_difference < responseQualityDifference: # 品質差が小さい場合(0.1未満) # → 軽量モデル(Haiku)を使用 # → "この程度の質問なら軽量モデルで十分"というケース use_lightweight_model() else: # 品質差が大きい場合(0.1以上) # → 高性能モデル(Sonnet)を使用 # → "この質問は高性能モデルを使う価値がある"というケース use_high_quality_model()
2. 仕組み 実行とフォールバック ルーティングで選択されたモデル(Sonnet or Haiku)で推論を行う ルーティング失敗時やタイムアウト時は、フォールバックモデル(Sonnet)を使 用して推論
3. 実際に試してみる 3 つのテストケースを用意: 1. シンプルな質問 2. 中程度の質問 3. 複雑な質問
テストケース 1: こんにちは 「こんにちは」
テストケース 1: こんにちは 「こんにちは」 → Sonnet が選択される
テストケース 2: EC2 について質問 「AWS の EC2 とは何ですか?一行で説明してください」
テストケース 2: EC2 について質問 「AWS の EC2 とは何ですか?一行で説明してください」 → Sonnet
が選択される
テストケース 3: 英語で質問 What is your name?
テストケース 3: 英語で質問 What is your name? → Haiku が選択される
テストケース 4: 英語で EC2 について質問 What is EC2?
テストケース 4: 英語で EC2 について質問 What is EC2? → Haiku
が選択される
考察 日本語で質問した場合、それだけでスコアが上がっている可能性がある そのため、閾値の調整が必要 英語の質問は適切にルーティングされてそう
5. まとめ Prompt Routing のメリット 1. 簡単にプロンプトの内容に応じたモデルを動的に使用できる 2. 適切なモデルを選ぶことによってコストの削減・応答時間の改善が期待できる 実装時の注意点
日本語の場合は閾値の調整が必要 プレビュー中なので閾値の調整はできない?
ご清聴ありがとうございました!