Ver3.0 ๐ค LangChain์ผ๋ก ๋ฉํฐ๋ชจ๋ฌ ๋ถ์ ๋ด ๋ง๋ค๊ธฐ

๐ค LangChain์ผ๋ก ๋ฉํฐ๋ชจ๋ฌ ๋ถ์ ๋ด ๋ง๋ค๊ธฐ
ํ ์คํธ, ์ด๋ฏธ์ง, ์ค๋์ค๋ฅผ ํ ๋ฒ์! ๋๋ํ AI ๋ด ์ ์ ์์ ์ ๋ณต
๐ฏ ๋ฉํฐ๋ชจ๋ฌ AI, ์ ์ง๊ธ ์ฃผ๋ชฉ๋ฐ์๊น?
์๋
, ์น๊ตฌ! ๐ ์์ฆ AI ๊ฐ๋ฐ ํธ๋ ๋๋ฅผ ๋ณด๋ฉด ์ ๋ง ์ ๊ธฐํ ๊ฒ ๋ง์. ์์ ์๋ ํ
์คํธ๋ง ์ดํดํ๋ ์ฑ๋ด์ด ๋๋ถ๋ถ์ด์๋๋ฐ, ์ด์ ๋ ์ฌ์ง์ ๋ณด์ฌ์ฃผ๋ฉด ๊ทธ ์์ ๋ญ๊ฐ ์๋์ง ์ค๋ช
ํด์ฃผ๊ณ , ์์ฑ์ ๋ค๋ ค์ฃผ๋ฉด ๊ฐ์ ๊น์ง ๋ถ์ํ๋ ์๋๊ฐ ์๊ฑฐ๋ !
์ด๊ฒ ๋ฐ๋ก ๋ฉํฐ๋ชจ๋ฌ(Multimodal) AI์ผ. ์ฌ๋ฌ ๊ฐ์ง ํํ์ ๋ฐ์ดํฐ๋ฅผ ๋์์ ์ฒ๋ฆฌํ ์ ์๋ ์ธ๊ณต์ง๋ฅ์ ๋งํ๋ ๊ฑฐ์ง. ํ
์คํธ, ์ด๋ฏธ์ง, ์ค๋์ค, ๋น๋์ค๊น์ง ํ๊บผ๋ฒ์ ์ดํดํ๊ณ ๋ถ์ํ ์ ์๋ค๋, ์ ๋ง ๋๋จํ์ง ์์? ๐
โข ์๋ฃ ๋ถ์ผ: X-ray ์ด๋ฏธ์ง์ ํ์ ์ฆ์ ํ ์คํธ๋ฅผ ๋์์ ๋ถ์ํด ์ง๋จ ๋ณด์กฐ
โข ์ผํ: ์ ํ ์ฌ์ง์ ์ฐ์ผ๋ฉด ์ ์ฌ ์ํ์ ์ฐพ์์ฃผ๊ณ ๋ฆฌ๋ทฐ๊น์ง ์์ฝ
โข ๊ต์ก: ํ์์ด ์ ์ถํ ๊ณผ์ (ํ ์คํธ+์ด๋ฏธ์ง)๋ฅผ ์ข ํฉ์ ์ผ๋ก ํ๊ฐ
โข ์ฝํ ์ธ ์ ์: ์์์ ์ฅ๋ฉด๊ณผ ๋์ฌ๋ฅผ ๋ถ์ํด ์๋์ผ๋ก ํ์ด๋ผ์ดํธ ์์ฑ
๊ทธ๋ฐ๋ฐ ์ด๋ฐ ๋ฉํฐ๋ชจ๋ฌ AI๋ฅผ ์ฒ์๋ถํฐ ๋ง๋ค๋ ค๋ฉด ์ ๋ง ๋ณต์กํด. ๊ฐ๊ฐ์ ๋ชจ๋ธ์ ๋ฐ๋ก ํ์ต์ํค๊ณ , ๊ฒฐ๊ณผ๋ฅผ ํตํฉํ๊ณ , ํ๋กฌํํธ๋ฅผ ๊ด๋ฆฌํ๊ณ ... ์๊ฐ๋ง ํด๋ ๋จธ๋ฆฌ๊ฐ ์ํ์ง? ๐ต
๋ฐ๋ก ์ฌ๊ธฐ์ LangChain์ด ๋ฑ์ฅํด! LangChain์ ์ด๋ฐ ๋ณต์กํ ๊ณผ์ ์ ํจ์ฌ ์ฝ๊ฒ ๋ง๋ค์ด์ฃผ๋ ํ๋ ์์ํฌ์ผ. ๋ง์น ๋ ๊ณ ๋ธ๋ก์ฒ๋ผ ํ์ํ ๊ธฐ๋ฅ๋ค์ ์กฐ๋ฆฝํด์ ๋๋ง์ AI ๋ด์ ๋ง๋ค ์ ์๊ฑฐ๋ . ๐งฉ
๐ง LangChain์ด ๋ญ๊ธธ๋?
LangChain์ 2022๋
์ ๋ฑ์ฅํ ์คํ์์ค ํ๋ ์์ํฌ์ผ. ๋๊ท๋ชจ ์ธ์ด ๋ชจ๋ธ(LLM)์ ํ์ฉํ ์ ํ๋ฆฌ์ผ์ด์
์ ์ฝ๊ฒ ๋ง๋ค ์ ์๋๋ก ๋์์ฃผ๋ ๋๊ตฌ์ธ๋ฐ, ํนํ ๋ฉํฐ๋ชจ๋ฌ ์ฒ๋ฆฌ์ ์ ๋ง ๊ฐ๋ ฅํ ๊ธฐ๋ฅ์ ์ ๊ณตํด. ๐ ๏ธ
์๊ฐํด๋ด. GPT-4, Claude, Gemini ๊ฐ์ ์ต์ AI ๋ชจ๋ธ๋ค์ ์ด๋ฏธ ํ
์คํธ์ ์ด๋ฏธ์ง๋ฅผ ๋์์ ์ฒ๋ฆฌํ ์ ์์์? ๊ทผ๋ฐ ์ด๊ฑธ ์ค์ ์๋น์ค์ ์ ์ฉํ๋ ค๋ฉด API ํธ์ถ, ๋ฐ์ดํฐ ์ ์ฒ๋ฆฌ, ๊ฒฐ๊ณผ ํ์ฒ๋ฆฌ, ์๋ฌ ํธ๋ค๋ง ๋ฑ๋ฑ... ์ ๊ฒฝ ์ธ ๊ฒ ํ๋ ๊ฐ์ง๊ฐ ์๋์ผ.
1. ๋ชจ๋ํ๋ ๊ตฌ์กฐ: ํ์ํ ์ปดํฌ๋ํธ๋ง ๊ณจ๋ผ์ ์ฌ์ฉ ๊ฐ๋ฅ
2. ๋ค์ํ ๋ชจ๋ธ ์ง์: OpenAI, Anthropic, Google, HuggingFace ๋ฑ ์ฃผ์ AI ๋ชจ๋ธ ํตํฉ
3. ์ฒด์ธ ๊ตฌ์ฑ: ์ฌ๋ฌ ์์ ์ ์์ฐจ์ ์ผ๋ก ์ฐ๊ฒฐํด ๋ณต์กํ ์ํฌํ๋ก์ฐ ๊ตฌํ
4. ๋ฉ๋ชจ๋ฆฌ ๊ด๋ฆฌ: ๋ํ ๋งฅ๋ฝ์ ๊ธฐ์ตํ๊ณ ์ ์งํ๋ ๊ธฐ๋ฅ
5. ์์ด์ ํธ ์์คํ : ์ค์ค๋ก ํ๋จํ๊ณ ๋๊ตฌ๋ฅผ ์ฌ์ฉํ๋ ์์จ AI ๊ตฌํ
ํนํ ๋ฉํฐ๋ชจ๋ฌ ๋ถ์ ๋ด์ ๋ง๋ค ๋ LangChain์ด ๋น์ ๋ฐํ๋ ์ด์ ๋ ๋ค์ํ ์ ๋ ฅ ํ์์ ์๋์ผ๋ก ์ฒ๋ฆฌํด์ฃผ๊ธฐ ๋๋ฌธ์ด์ผ. ์๋ฅผ ๋ค์ด, ์ฌ์ฉ์๊ฐ ์ด๋ฏธ์ง์ ํ ์คํธ๋ฅผ ํจ๊ป ๋ณด๋ด๋ฉด LangChain์ด ์์์ ์ ์ ํ ํ์์ผ๋ก ๋ณํํด์ AI ๋ชจ๋ธ์ ์ ๋ฌํด์ฃผ๊ฑฐ๋ . ๐
๐ ๊ฐ๋ฐ ํ๊ฒฝ ์ค๋นํ๊ธฐ
์, ์ด์ ๋ณธ๊ฒฉ์ ์ผ๋ก ๋ฉํฐ๋ชจ๋ฌ ๋ถ์ ๋ด์ ๋ง๋ค์ด๋ณผ๊น? ๋จผ์ ๊ฐ๋ฐ ํ๊ฒฝ๋ถํฐ ์ธํ ํด์ผ ํด. ๊ฑฑ์ ๋ง, ์๊ฐ๋ณด๋ค ์ด๋ ต์ง ์์! ์ฐจ๊ทผ์ฐจ๊ทผ ๋ฐ๋ผ์ค๋ฉด ๋ผ. โ
LangChain์ Python ๊ธฐ๋ฐ์ด์ผ. Python 3.8 ์ด์ ๋ฒ์ ์ด ํ์ํด. ํฐ๋ฏธ๋์์ ํ์ธํด๋ณด์:
python --version
# ๋๋
python3 --version
๋ง์ฝ ์ค์น๋์ด ์์ง ์๋ค๋ฉด python.org์์ ๋ค์ด๋ก๋ํด์ ์ค์นํ๋ฉด ๋ผ.
ํ๋ก์ ํธ๋ณ๋ก ๋ ๋ฆฝ์ ์ธ ํ๊ฒฝ์ ๋ง๋๋ ๊ฒ ์ข์. ํจํค์ง ์ถฉ๋์ ๋ฐฉ์งํ ์ ์๊ฑฐ๋ !
# ๊ฐ์ํ๊ฒฝ ์์ฑ
python -m venv multimodal_bot
# ๊ฐ์ํ๊ฒฝ ํ์ฑํ (Windows)
multimodal_bot\Scripts\activate
# ๊ฐ์ํ๊ฒฝ ํ์ฑํ (Mac/Linux)
source multimodal_bot/bin/activate
์ด์ ํ์ํ ๋ผ์ด๋ธ๋ฌ๋ฆฌ๋ค์ ์ค์นํ ์ฐจ๋ก์ผ. ํ ๋ฒ์ ์ค์นํ๋ฉด ํธํด:
pip install langchain langchain-openai langchain-anthropic
pip install pillow python-dotenv
pip install openai anthropic
pip install langchain-community
๊ฐ ํจํค์ง์ ์ญํ ์ ๊ฐ๋จํ ์ค๋ช
ํ๋ฉด:โข langchain: ํต์ฌ ํ๋ ์์ํฌ
โข langchain-openai: OpenAI ๋ชจ๋ธ ์ฐ๋
โข pillow: ์ด๋ฏธ์ง ์ฒ๋ฆฌ
โข python-dotenv: ํ๊ฒฝ๋ณ์ ๊ด๋ฆฌ
AI ๋ชจ๋ธ์ ์ฌ์ฉํ๋ ค๋ฉด API ํค๊ฐ ํ์ํด. OpenAI๋ Anthropic ์น์ฌ์ดํธ์์ ๋ฐ๊ธ๋ฐ์ ์ ์์ด.
ํ๋ก์ ํธ ํด๋์ .env ํ์ผ์ ๋ง๋ค๊ณ :
OPENAI_API_KEY=your_openai_api_key_here
ANTHROPIC_API_KEY=your_anthropic_api_key_here
โ ๏ธ ์ฃผ์: API ํค๋ ์ ๋ GitHub ๊ฐ์ ๊ณต๊ฐ ์ ์ฅ์์ ์ฌ๋ฆฌ๋ฉด ์ ๋ผ! .gitignore ํ์ผ์ .env๋ฅผ ์ถ๊ฐํด๋์.
๐ป ๊ธฐ๋ณธ ๋ฉํฐ๋ชจ๋ฌ ๋ด ๊ตฌํํ๊ธฐ
ํ๊ฒฝ ์ค์ ์ด ๋๋ฌ์ผ๋ ์ด์ ์ง์ง ์ฝ๋ฉ์ ์์ํด๋ณผ๊น? ๋จผ์ ๊ฐ๋จํ ํ ์คํธ+์ด๋ฏธ์ง ๋ถ์ ๋ด๋ถํฐ ๋ง๋ค์ด๋ณด์. ์ด๊ฒ ๊ธฐ๋ณธ์ด ๋๋ฉด ๋์ค์ ์ค๋์ค๋ ๋น๋์ค๋ ์ถ๊ฐํ ์ ์์ด! ๐จ
๐ 1๋จ๊ณ: ๊ธฐ๋ณธ ๊ตฌ์กฐ ๋ง๋ค๊ธฐ
๋จผ์ ํ๋ก์ ํธ์ ๋ผ๋๋ฅผ ๋ง๋ค์ด์ผ ํด. multimodal_bot.py ํ์ผ์ ์์ฑํ๊ณ ๋ค์ ์ฝ๋๋ฅผ ์์ฑํด๋ณด์:
import os
from dotenv import load_dotenv
from langchain_openai import ChatOpenAI
from langchain.schema.messages import HumanMessage, SystemMessage
from langchain.prompts import ChatPromptTemplate
import base64
from pathlib import Path
# ํ๊ฒฝ๋ณ์ ๋ก๋
load_dotenv()
class MultimodalBot:
def __init__(self, model_name="gpt-4-vision-preview"):
"""
๋ฉํฐ๋ชจ๋ฌ ๋ถ์ ๋ด ์ด๊ธฐํ
Args:
model_name: ์ฌ์ฉํ AI ๋ชจ๋ธ ์ด๋ฆ
"""
self.llm = ChatOpenAI(
model=model_name,
temperature=0.7,
max_tokens=1000
)
def encode_image(self, image_path):
"""
์ด๋ฏธ์ง๋ฅผ base64๋ก ์ธ์ฝ๋ฉ
Args:
image_path: ์ด๋ฏธ์ง ํ์ผ ๊ฒฝ๋ก
Returns:
base64๋ก ์ธ์ฝ๋ฉ๋ ์ด๋ฏธ์ง ๋ฌธ์์ด
"""
with open(image_path, "rb") as image_file:
return base64.b64encode(image_file.read()).decode('utf-8')
def analyze_image_with_text(self, image_path, user_question):
"""
์ด๋ฏธ์ง์ ํ
์คํธ๋ฅผ ํจ๊ป ๋ถ์
Args:
image_path: ๋ถ์ํ ์ด๋ฏธ์ง ๊ฒฝ๋ก
user_question: ์ฌ์ฉ์ ์ง๋ฌธ
Returns:
AI์ ๋ถ์ ๊ฒฐ๊ณผ
"""
# ์ด๋ฏธ์ง ์ธ์ฝ๋ฉ
base64_image = self.encode_image(image_path)
# ๋ฉ์์ง ๊ตฌ์ฑ
messages = [
SystemMessage(content="๋น์ ์ ์ด๋ฏธ์ง๋ฅผ ๋ถ์ํ๊ณ ์ฌ์ฉ์์ ์ง๋ฌธ์ ๋ต๋ณํ๋ ์ ๋ฌธ๊ฐ์
๋๋ค."),
HumanMessage(
content=[
{"type": "text", "text": user_question},
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}"
}
}
]
)
]
# AI ๋ชจ๋ธ ํธ์ถ
response = self.llm.invoke(messages)
return response.content
# ์ฌ์ฉ ์์
if __name__ == "__main__":
bot = MultimodalBot()
# ์ด๋ฏธ์ง ๊ฒฝ๋ก์ ์ง๋ฌธ ์ค์
image_path = "sample_image.jpg"
question = "์ด ์ด๋ฏธ์ง์์ ๋ฌด์์ ๋ณผ ์ ์๋์? ์์ธํ ์ค๋ช
ํด์ฃผ์ธ์."
# ๋ถ์ ์คํ
result = bot.analyze_image_with_text(image_path, question)
print(f"๋ถ์ ๊ฒฐ๊ณผ:\n{result}")
1. ํด๋์ค ๊ตฌ์กฐ: MultimodalBot ํด๋์ค๋ก ๋ชจ๋ ๊ธฐ๋ฅ์ ์บก์ํํ์ด. ์ฌ์ฌ์ฉํ๊ธฐ ํธํ์ง?
2. ์ด๋ฏธ์ง ์ธ์ฝ๋ฉ: AI ๋ชจ๋ธ์ ์ด๋ฏธ์ง๋ฅผ ์ง์ ๋ฐ์ ์ ์์ด์ base64 ๋ฌธ์์ด๋ก ๋ณํํด์ผ ํด.
3. ๋ฉ์์ง ๊ตฌ์ฑ: LangChain์ ๋ฉ์์ง ์์คํ ์ ์ฌ์ฉํด์ ํ ์คํธ์ ์ด๋ฏธ์ง๋ฅผ ํจ๊ป ์ ๋ฌํด.
4. ์ ์ฐํ ์ค๊ณ: ๋์ค์ ๊ธฐ๋ฅ์ ์ถ๊ฐํ๊ธฐ ์ฝ๋๋ก ๋ฉ์๋๋ฅผ ๋ถ๋ฆฌํ์ด.
๐ฏ 2๋จ๊ณ: ๊ณ ๊ธ ๊ธฐ๋ฅ ์ถ๊ฐํ๊ธฐ
๊ธฐ๋ณธ ๋ด์ด ์๋ํ๋ ๊ฑธ ํ์ธํ๋ค๋ฉด, ์ด์ ๋ ๋๋ํ๊ฒ ๋ง๋ค์ด๋ณด์! ์ฌ๋ฌ ์ด๋ฏธ์ง๋ฅผ ๋์์ ๋ถ์ํ๊ฑฐ๋, ๋ํ ๋งฅ๋ฝ์ ๊ธฐ์ตํ๋ ๊ธฐ๋ฅ์ ์ถ๊ฐํ ์ ์์ด. ๐ง
from langchain.memory import ConversationBufferMemory
from langchain.chains import ConversationChain
from typing import List, Dict
class AdvancedMultimodalBot(MultimodalBot):
def __init__(self, model_name="gpt-4-vision-preview"):
super().__init__(model_name)
# ๋ํ ๊ธฐ๋ก์ ์ ์ฅํ ๋ฉ๋ชจ๋ฆฌ
self.memory = ConversationBufferMemory(
return_messages=True,
memory_key="chat_history"
)
def analyze_multiple_images(self, image_paths: List[str], question: str):
"""
์ฌ๋ฌ ์ด๋ฏธ์ง๋ฅผ ๋์์ ๋ถ์
Args:
image_paths: ์ด๋ฏธ์ง ํ์ผ ๊ฒฝ๋ก ๋ฆฌ์คํธ
question: ์ฌ์ฉ์ ์ง๋ฌธ
Returns:
์ข
ํฉ ๋ถ์ ๊ฒฐ๊ณผ
"""
# ์ด๋ฏธ์ง๋ค์ ์ธ์ฝ๋ฉ
encoded_images = [self.encode_image(path) for path in image_paths]
# ๋ฉ์์ง ์ปจํ
์ธ ๊ตฌ์ฑ
content = [{"type": "text", "text": question}]
for idx, encoded_img in enumerate(encoded_images):
content.append({
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{encoded_img}",
"detail": "high" # ๊ณ ํด์๋ ๋ถ์
}
})
messages = [
SystemMessage(content="์ฌ๋ฌ ์ด๋ฏธ์ง๋ฅผ ๋น๊ต ๋ถ์ํ๋ ์ ๋ฌธ๊ฐ์
๋๋ค."),
HumanMessage(content=content)
]
response = self.llm.invoke(messages)
return response.content
def analyze_with_context(self, image_path: str, question: str):
"""
์ด์ ๋ํ ๋งฅ๋ฝ์ ๊ณ ๋ คํ ๋ถ์
Args:
image_path: ์ด๋ฏธ์ง ๊ฒฝ๋ก
question: ์ฌ์ฉ์ ์ง๋ฌธ
Returns:
๋งฅ๋ฝ์ ๊ณ ๋ คํ ๋ต๋ณ
"""
# ์ด์ ๋ํ ๊ธฐ๋ก ๊ฐ์ ธ์ค๊ธฐ
chat_history = self.memory.load_memory_variables({})
base64_image = self.encode_image(image_path)
# ์์คํ
ํ๋กฌํํธ์ ๋งฅ๋ฝ ์ ๋ณด ์ถ๊ฐ
context_prompt = "์ด์ ๋ํ ๋ด์ฉ์ ์ฐธ๊ณ ํ์ฌ ๋ต๋ณํด์ฃผ์ธ์."
if chat_history.get("chat_history"):
context_prompt += f"\n์ด์ ๋ํ: {chat_history['chat_history']}"
messages = [
SystemMessage(content=context_prompt),
HumanMessage(
content=[
{"type": "text", "text": question},
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}"
}
}
]
)
]
response = self.llm.invoke(messages)
# ๋ํ ๊ธฐ๋ก ์ ์ฅ
self.memory.save_context(
{"input": question},
{"output": response.content}
)
return response.content
def extract_structured_data(self, image_path: str, schema: Dict):
"""
์ด๋ฏธ์ง์์ ๊ตฌ์กฐํ๋ ๋ฐ์ดํฐ ์ถ์ถ
Args:
image_path: ์ด๋ฏธ์ง ๊ฒฝ๋ก
schema: ์ถ์ถํ ๋ฐ์ดํฐ ์คํค๋ง
Returns:
๊ตฌ์กฐํ๋ ๋ฐ์ดํฐ
"""
base64_image = self.encode_image(image_path)
# ์คํค๋ง๋ฅผ ํ๋กฌํํธ๋ก ๋ณํ
schema_text = "๋ค์ ํ์์ผ๋ก ์ ๋ณด๋ฅผ ์ถ์ถํด์ฃผ์ธ์:\n"
for key, description in schema.items():
schema_text += f"- {key}: {description}\n"
messages = [
SystemMessage(content="์ด๋ฏธ์ง์์ ์ ๋ณด๋ฅผ ์ถ์ถํ์ฌ JSON ํ์์ผ๋ก ๋ฐํํฉ๋๋ค."),
HumanMessage(
content=[
{"type": "text", "text": schema_text},
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}"
}
}
]
)
]
response = self.llm.invoke(messages)
return response.content
# ๊ณ ๊ธ ๊ธฐ๋ฅ ์ฌ์ฉ ์์
if __name__ == "__main__":
bot = AdvancedMultimodalBot()
# ์์ 1: ์ฌ๋ฌ ์ด๋ฏธ์ง ๋น๊ต ๋ถ์
images = ["image1.jpg", "image2.jpg", "image3.jpg"]
comparison = bot.analyze_multiple_images(
images,
"์ด ์ธ ์ด๋ฏธ์ง์ ๊ณตํต์ ๊ณผ ์ฐจ์ด์ ์ ๋ถ์ํด์ฃผ์ธ์."
)
print(f"๋น๊ต ๋ถ์:\n{comparison}\n")
# ์์ 2: ๋งฅ๋ฝ ๊ธฐ๋ฐ ๋ํ
result1 = bot.analyze_with_context(
"product.jpg",
"์ด ์ ํ์ ํน์ง์ ์ค๋ช
ํด์ฃผ์ธ์."
)
print(f"์ฒซ ๋ฒ์งธ ๋ต๋ณ:\n{result1}\n")
result2 = bot.analyze_with_context(
"product.jpg",
"๊ทธ๋ ๋ค๋ฉด ์ด ์ ํ์ ๊ฐ๊ฒฉ๋๋ ์ด๋ ์ ๋์ผ๊น์?"
)
print(f"๋ ๋ฒ์งธ ๋ต๋ณ (๋งฅ๋ฝ ๊ณ ๋ ค):\n{result2}\n")
# ์์ 3: ๊ตฌ์กฐํ๋ ๋ฐ์ดํฐ ์ถ์ถ
schema = {
"์ ํ๋ช
": "์ด๋ฏธ์ง์ ํ์๋ ์ ํ์ ์ด๋ฆ",
"์์": "์ ํ์ ์ฃผ์ ์์",
"ํน์ง": "๋์ ๋๋ ํน์ง๋ค",
"์ฉ๋": "์์๋๋ ์ฌ์ฉ ์ฉ๋"
}
structured_data = bot.extract_structured_data("product.jpg", schema)
print(f"์ถ์ถ๋ ๋ฐ์ดํฐ:\n{structured_data}")
๋ฉ๋ชจ๋ฆฌ ๊ด๋ฆฌ: ๋ํ๊ฐ ๊ธธ์ด์ง๋ฉด ๋ฉ๋ชจ๋ฆฌ๊ฐ ๊ณ์ ์์ฌ. ์ฃผ๊ธฐ์ ์ผ๋ก memory.clear()๋ฅผ ํธ์ถํด์ ์ด๊ธฐํํ๋ ๊ฒ ์ข์.
์ด๋ฏธ์ง ํฌ๊ธฐ: ๋๋ฌด ํฐ ์ด๋ฏธ์ง๋ API ๋น์ฉ์ด ๋ง์ด ๋ค์ด. ์ ์ ํ ๋ฆฌ์ฌ์ด์งํ๋ ๊ฒ ๊ฒฝ์ ์ ์ด์ผ.
์๋ฌ ์ฒ๋ฆฌ: ์ค์ ์๋น์ค์์๋ try-except ๋ธ๋ก์ผ๋ก ์์ธ ์ฒ๋ฆฌ๋ฅผ ๊ผญ ํด์ผ ํด!
๐ต ์ค๋์ค ๋ถ์ ๊ธฐ๋ฅ ์ถ๊ฐํ๊ธฐ
์ด๋ฏธ์ง ๋ถ์์ ๋ง์คํฐํ์ผ๋, ์ด์ ์ค๋์ค๋ ๋ค๋ค๋ณผ๊น? ์์ฑ ํ์ผ์ ํ
์คํธ๋ก ๋ณํํ๊ณ , ๊ฐ์ ์ ๋ถ์ํ๊ณ , ๋ด์ฉ์ ์์ฝํ๋ ๊ธฐ๋ฅ์ ๋ง๋ค์ด๋ณด์! ๐ง
์ค๋์ค ์ฒ๋ฆฌ๋ฅผ ์ํด์๋ ์ถ๊ฐ ๋ผ์ด๋ธ๋ฌ๋ฆฌ๊ฐ ํ์ํด:
pip install openai-whisper pydub
# ๋๋ OpenAI API๋ฅผ ์ฌ์ฉํ๋ ๊ฒฝ์ฐ
pip install openai
Whisper๋ OpenAI๊ฐ ๋ง๋ ์์ฑ ์ธ์ ๋ชจ๋ธ์ด์ผ. ์ ํ๋๊ฐ ์ ๋ง ๋ฐ์ด๋๊ณ , ๋ค์ํ ์ธ์ด๋ฅผ ์ง์ํด. ๐
import whisper
from pydub import AudioSegment
import tempfile
import os
class AudioMultimodalBot(AdvancedMultimodalBot):
def __init__(self, model_name="gpt-4-vision-preview", whisper_model="base"):
super().__init__(model_name)
# Whisper ๋ชจ๋ธ ๋ก๋
self.whisper = whisper.load_model(whisper_model)
def transcribe_audio(self, audio_path: str):
"""
์ค๋์ค ํ์ผ์ ํ
์คํธ๋ก ๋ณํ
Args:
audio_path: ์ค๋์ค ํ์ผ ๊ฒฝ๋ก
Returns:
๋ณํ๋ ํ
์คํธ
"""
result = self.whisper.transcribe(audio_path)
return result["text"]
def analyze_audio_sentiment(self, audio_path: str):
"""
์ค๋์ค์ ๊ฐ์ ๋ถ์
Args:
audio_path: ์ค๋์ค ํ์ผ ๊ฒฝ๋ก
Returns:
๊ฐ์ ๋ถ์ ๊ฒฐ๊ณผ
"""
# ๋จผ์ ํ
์คํธ๋ก ๋ณํ
transcript = self.transcribe_audio(audio_path)
# ๊ฐ์ ๋ถ์ ํ๋กฌํํธ
messages = [
SystemMessage(content="์์ฑ ๋ด์ฉ์ ๊ฐ์ ์ ๋ถ์ํ๋ ์ ๋ฌธ๊ฐ์
๋๋ค."),
HumanMessage(content=f"""
๋ค์ ์์ฑ ๋ด์ฉ์ ๊ฐ์ ์ ๋ถ์ํด์ฃผ์ธ์:
"{transcript}"
๋ค์ ํญ๋ชฉ์ ํฌํจํด์ฃผ์ธ์:
1. ์ ๋ฐ์ ์ธ ๊ฐ์ (๊ธ์ /๋ถ์ /์ค๋ฆฝ)
2. ๊ตฌ์ฒด์ ์ธ ๊ฐ์ (๊ธฐ์จ, ์ฌํ, ๋ถ๋
ธ, ๋๋ ๋ฑ)
3. ๊ฐ์ ์ ๊ฐ๋ (1-10)
4. ์ฃผ์ ํค์๋
""")
]
response = self.llm.invoke(messages)
return {
"transcript": transcript,
"sentiment_analysis": response.content
}
def summarize_audio(self, audio_path: str, summary_length="medium"):
"""
์ค๋์ค ๋ด์ฉ ์์ฝ
Args:
audio_path: ์ค๋์ค ํ์ผ ๊ฒฝ๋ก
summary_length: ์์ฝ ๊ธธ์ด (short/medium/long)
Returns:
์์ฝ๋ ๋ด์ฉ
"""
transcript = self.transcribe_audio(audio_path)
length_instructions = {
"short": "3-5๋ฌธ์ฅ์ผ๋ก ํต์ฌ๋ง ๊ฐ๋จํ",
"medium": "1-2 ๋จ๋ฝ์ผ๋ก ์ฃผ์ ๋ด์ฉ์",
"long": "์์ธํ๊ฒ ๋ชจ๋ ์ค์ ํฌ์ธํธ๋ฅผ"
}
messages = [
SystemMessage(content="์์ฑ ๋ด์ฉ์ ์์ฝํ๋ ์ ๋ฌธ๊ฐ์
๋๋ค."),
HumanMessage(content=f"""
๋ค์ ์์ฑ ๋ด์ฉ์ {length_instructions.get(summary_length, length_instructions['medium'])} ์์ฝํด์ฃผ์ธ์:
"{transcript}"
""")
]
response = self.llm.invoke(messages)
return {
"original_transcript": transcript,
"summary": response.content
}
def multimodal_analysis(self, image_path: str, audio_path: str, question: str):
"""
์ด๋ฏธ์ง์ ์ค๋์ค๋ฅผ ํจ๊ป ๋ถ์
Args:
image_path: ์ด๋ฏธ์ง ํ์ผ ๊ฒฝ๋ก
audio_path: ์ค๋์ค ํ์ผ ๊ฒฝ๋ก
question: ์ฌ์ฉ์ ์ง๋ฌธ
Returns:
ํตํฉ ๋ถ์ ๊ฒฐ๊ณผ
"""
# ์ค๋์ค๋ฅผ ํ
์คํธ๋ก ๋ณํ
audio_transcript = self.transcribe_audio(audio_path)
# ์ด๋ฏธ์ง ์ธ์ฝ๋ฉ
base64_image = self.encode_image(image_path)
# ํตํฉ ๋ถ์ ๋ฉ์์ง ๊ตฌ์ฑ
messages = [
SystemMessage(content="์ด๋ฏธ์ง์ ์์ฑ ๋ด์ฉ์ ์ข
ํฉ์ ์ผ๋ก ๋ถ์ํ๋ ์ ๋ฌธ๊ฐ์
๋๋ค."),
HumanMessage(
content=[
{
"type": "text",
"text": f"""
์ฌ์ฉ์ ์ง๋ฌธ: {question}
์์ฑ ๋ด์ฉ: "{audio_transcript}"
์ ์์ฑ ๋ด์ฉ๊ณผ ์๋ ์ด๋ฏธ์ง๋ฅผ ํจ๊ป ๊ณ ๋ คํ์ฌ ๋ต๋ณํด์ฃผ์ธ์.
"""
},
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}"
}
}
]
)
]
response = self.llm.invoke(messages)
return {
"audio_transcript": audio_transcript,
"analysis": response.content
}
# ์ค๋์ค ๊ธฐ๋ฅ ์ฌ์ฉ ์์
if __name__ == "__main__":
bot = AudioMultimodalBot()
# ์์ 1: ์์ฑ ๊ฐ์ ๋ถ์
sentiment = bot.analyze_audio_sentiment("speech.mp3")
print(f"์์ฑ ๋ด์ฉ: {sentiment['transcript']}")
print(f"๊ฐ์ ๋ถ์: {sentiment['sentiment_analysis']}\n")
# ์์ 2: ์์ฑ ์์ฝ
summary = bot.summarize_audio("lecture.mp3", summary_length="medium")
print(f"์์ฝ: {summary['summary']}\n")
# ์์ 3: ์ด๋ฏธ์ง + ์ค๋์ค ํตํฉ ๋ถ์
result = bot.multimodal_analysis(
"presentation_slide.jpg",
"presentation_audio.mp3",
"์ด ํ๋ ์ ํ
์ด์
์ ํต์ฌ ๋ฉ์์ง๋ ๋ฌด์์ธ๊ฐ์?"
)
print(f"ํตํฉ ๋ถ์: {result['analysis']}")
Whisper ๋ชจ๋ธ ํฌ๊ธฐ: "base" ๋ชจ๋ธ์ ๋น ๋ฅด์ง๋ง ์ ํ๋๊ฐ ๋ฎ์. "medium"์ด๋ "large" ๋ชจ๋ธ์ ์ฌ์ฉํ๋ฉด ๋ ์ ํํ์ง๋ง ์ฒ๋ฆฌ ์๊ฐ์ด ๊ธธ์ด์ ธ.
์ค๋์ค ํ์: Whisper๋ ๋๋ถ๋ถ์ ์ค๋์ค ํ์์ ์ง์ํ์ง๋ง, WAV๋ MP3๊ฐ ๊ฐ์ฅ ์์ ์ ์ด์ผ.
ํ์ผ ํฌ๊ธฐ: ๊ธด ์ค๋์ค๋ ์ฒ๋ฆฌ ์๊ฐ์ด ์ค๋ ๊ฑธ๋ ค. ํ์ํ๋ฉด ์ฒญํฌ๋ก ๋๋ ์ ์ฒ๋ฆฌํ๋ ๊ฒ ์ข์.
๐ ๏ธ ์ค์ ํ์ฉ ์ฌ๋ก
์ด๋ก ๊ณผ ์ฝ๋๋ ์ถฉ๋ถํ ๋ดค์ผ๋, ์ด์ ์ค์ ๋ก ์ด๋ป๊ฒ ํ์ฉํ ์ ์๋์ง ๊ตฌ์ฒด์ ์ธ ์์๋ฅผ ์ดํด๋ณด์! ๋ค์ํ ์ฐ์ ๋ถ์ผ์์ ๋ฉํฐ๋ชจ๋ฌ ๋ด์ด ์ด๋ป๊ฒ ์ฐ์ด๋์ง ์์๋ณด๋ฉด ์์ด๋์ด๊ฐ ์์์ ๊ฑฐ์ผ. ๐ก
๐ฅ ์๋ฃ ๋ถ์ผ: ์๋ฃ ์์ ๋ถ์ ๋ณด์กฐ
์์ฌ๊ฐ X-ray๋ CT ์ค์บ ์ด๋ฏธ์ง๋ฅผ ์ ๋ก๋ํ๊ณ , ํ์์ ์ฆ์์ ํ ์คํธ๋ก ์ ๋ ฅํ๋ฉด AI๊ฐ ์ด๊ธฐ ๋ถ์์ ์ ๊ณตํด. ๋ฌผ๋ก ์ต์ข ์ง๋จ์ ์์ฌ๊ฐ ํ์ง๋ง, ๋์น ์ ์๋ ๋ถ๋ถ์ ์ฒดํฌํ๋ ๋ฐ ๋์์ด ๋ผ.
class MedicalAnalysisBot(AudioMultimodalBot):
def analyze_medical_image(self, image_path: str, patient_info: dict):
"""
์๋ฃ ์์ ๋ถ์
Args:
image_path: ์๋ฃ ์์ ๊ฒฝ๋ก
patient_info: ํ์ ์ ๋ณด (๋์ด, ์ฑ๋ณ, ์ฆ์ ๋ฑ)
Returns:
๋ถ์ ๊ฒฐ๊ณผ ๋ฐ ์ฃผ์์ฌํญ
"""
base64_image = self.encode_image(image_path)
patient_context = f"""
ํ์ ์ ๋ณด:
- ๋์ด: {patient_info.get('age', '๋ฏธ์')}
- ์ฑ๋ณ: {patient_info.get('gender', '๋ฏธ์')}
- ์ฃผ์ ์ฆ์: {patient_info.get('symptoms', '๋ฏธ์')}
- ๋ณ๋ ฅ: {patient_info.get('history', '์์')}
"""
messages = [
SystemMessage(content="""
๋น์ ์ ์๋ฃ ์์ ๋ถ์์ ๋ณด์กฐํ๋ AI์
๋๋ค.
์ฃผ์: ์ด ๋ถ์์ ์ฐธ๊ณ ์ฉ์ด๋ฉฐ, ์ต์ข
์ง๋จ์ ๋ฐ๋์ ์๋ฃ ์ ๋ฌธ๊ฐ๊ฐ ํด์ผ ํฉ๋๋ค.
"""),
HumanMessage(
content=[
{"type": "text", "text": f"""
{patient_context}
์ ํ์์ ์๋ฃ ์์์ ๋ถ์ํด์ฃผ์ธ์:
1. ๊ด์ฐฐ๋๋ ์ฃผ์ ์๊ฒฌ
2. ์ฃผ์๊ฐ ํ์ํ ๋ถ๋ถ
3. ์ถ๊ฐ ๊ฒ์ฌ ๊ถ์ฅ์ฌํญ
4. ์ผ๋ฐ์ ์ธ ๊ฐ๋ณ ์ง๋จ ๋ชฉ๋ก
"""},
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}"
}
}
]
)
]
response = self.llm.invoke(messages)
return response.content
# ์ฌ์ฉ ์์
medical_bot = MedicalAnalysisBot()
patient = {
'age': 45,
'gender': '๋จ์ฑ',
'symptoms': 'ํํต, ํธํก๊ณค๋',
'history': '๊ณ ํ์'
}
analysis = medical_bot.analyze_medical_image("chest_xray.jpg", patient)
print(analysis)
๐๏ธ ์ด์ปค๋จธ์ค: ์ค๋งํธ ์ํ ์ถ์ฒ
๊ณ ๊ฐ์ด ์ํ๋ ์คํ์ผ์ ์ท ์ฌ์ง์ ์ฐ์ด์ ์ฌ๋ฆฌ๊ณ , "์ด๋ฐ ๋๋์ ์ท์ ์ฐพ๊ณ ์์ด์"๋ผ๊ณ ๋งํ๋ฉด AI๊ฐ ๋น์ทํ ์ํ์ ์ฐพ์์ฃผ๊ณ ์ฝ๋ ํ๊น์ง ์ ๊ณตํด. ์ฌ๋ฅ๋ท ๊ฐ์ ํ๋ซํผ์์๋ ์ด๋ฐ ๊ธฐ์ ๋ก ์ฌ์ฉ์์๊ฒ ๋ฑ ๋ง๋ ์ฌ๋ฅ์ ์ถ์ฒํ ์ ์๊ฒ ์ง? ๐จ
class EcommerceBot(AudioMultimodalBot):
def __init__(self, model_name="gpt-4-vision-preview"):
super().__init__(model_name)
# ์ํ ๋ฐ์ดํฐ๋ฒ ์ด์ค (์ค์ ๋ก๋ DB ์ฐ๋)
self.product_db = []
def analyze_style_preference(self, reference_images: List[str],
user_description: str):
"""
์ฌ์ฉ์์ ์คํ์ผ ์ ํธ๋ ๋ถ์
Args:
reference_images: ์ฐธ๊ณ ์ด๋ฏธ์ง๋ค
user_description: ์ฌ์ฉ์ ์ค๋ช
Returns:
์คํ์ผ ํ๋กํ ๋ฐ ์ถ์ฒ
"""
encoded_images = [self.encode_image(img) for img in reference_images]
content = [
{"type": "text", "text": f"""
์ฌ์ฉ์ ์ค๋ช
: "{user_description}"
์ ์ค๋ช
๊ณผ ์๋ ์ด๋ฏธ์ง๋ค์ ๋ถ์ํ์ฌ:
1. ์ ํธํ๋ ์คํ์ผ (์บ์ฃผ์ผ, ํฌ๋ฉ, ์คํธ๋ฆฟ ๋ฑ)
2. ์ ํธ ์์ ํ๋ ํธ
3. ์ฃผ์ ๋์์ธ ์์
4. ์ถ์ฒ ์์ดํ
์นดํ
๊ณ ๋ฆฌ
5. ์ฝ๋ ํ
์ ์ ๊ณตํด์ฃผ์ธ์.
"""}
]
for encoded_img in encoded_images:
content.append({
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{encoded_img}"
}
})
messages = [
SystemMessage(content="ํจ์
์คํ์ผ๋ฆฌ์คํธ AI์
๋๋ค."),
HumanMessage(content=content)
]
response = self.llm.invoke(messages)
return response.content
def compare_products(self, product_images: List[str], criteria: str):
"""
์ฌ๋ฌ ์ํ ๋น๊ต
Args:
product_images: ๋น๊ตํ ์ํ ์ด๋ฏธ์ง๋ค
criteria: ๋น๊ต ๊ธฐ์ค
Returns:
์์ธ ๋น๊ต ๋ถ์
"""
encoded_images = [self.encode_image(img) for img in product_images]
content = [
{"type": "text", "text": f"""
๋ค์ ๊ธฐ์ค์ผ๋ก ์ํ๋ค์ ๋น๊ตํด์ฃผ์ธ์: {criteria}
๊ฐ ์ํ์:
1. ์ฅ๋จ์
2. ๊ฐ๊ฒฉ ๋๋น ๊ฐ์น (์ด๋ฏธ์ง์์ ํ์ง ์ถ์ )
3. ์ด์ธ๋ฆฌ๋ ์ํฉ/์คํ์ผ
4. ์ต์ข
์ถ์ฒ ์์
๋ฅผ ์ ๊ณตํด์ฃผ์ธ์.
"""}
]
for idx, encoded_img in enumerate(encoded_images, 1):
content.append({
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{encoded_img}"
}
})
messages = [
SystemMessage(content="์ํ ๋น๊ต ์ ๋ฌธ๊ฐ AI์
๋๋ค."),
HumanMessage(content=content)
]
response = self.llm.invoke(messages)
return response.content
# ์ฌ์ฉ ์์
ecommerce_bot = EcommerceBot()
# ์คํ์ผ ๋ถ์
style_analysis = ecommerce_bot.analyze_style_preference(
["style1.jpg", "style2.jpg"],
"ํธ์ํ๋ฉด์๋ ์ธ๋ จ๋ ๋๋์ ์ท์ ์ข์ํด์"
)
print(f"์คํ์ผ ๋ถ์:\n{style_analysis}\n")
# ์ํ ๋น๊ต
comparison = ecommerce_bot.compare_products(
["product1.jpg", "product2.jpg", "product3.jpg"],
"๋์์ธ, ํ์ง, ํ์ฉ๋"
)
print(f"์ํ ๋น๊ต:\n{comparison}")
๐ ๊ต์ก: ํ์ต ์๋ฃ ๋ถ์ ๋ฐ ํผ๋๋ฐฑ
ํ์์ด ์์ผ๋ก ์ด ์ํ ๋ฌธ์ ํ์ด๋ฅผ ์ฌ์ง์ผ๋ก ์ฐ์ด ์ฌ๋ฆฌ๊ณ , ํ์ด ๊ณผ์ ์ ์ค๋ช ํ๋ ์์ฑ์ ๋ น์ํ๋ฉด AI๊ฐ ์ ๋ต ์ฌ๋ถ๋ฅผ ํ์ธํ๊ณ ์์ธํ ํผ๋๋ฐฑ์ ์ ๊ณตํด. ํ๋ฆฐ ๋ถ๋ถ์ด ์์ผ๋ฉด ์ด๋์ ์ค์ํ๋์ง ์ ํํ ์ง์ด์ค.
class EducationBot(AudioMultimodalBot):
def evaluate_homework(self, image_path: str, audio_path: str = None,
subject: str = "์ํ"):
"""
๊ณผ์ ํ๊ฐ ๋ฐ ํผ๋๋ฐฑ
Args:
image_path: ๊ณผ์ ์ด๋ฏธ์ง (์๊ธ์จ, ๊ทธ๋ฆผ ๋ฑ)
audio_path: ํ์์ ์ค๋ช
์์ฑ (์ ํ)
subject: ๊ณผ๋ชฉ
Returns:
ํ๊ฐ ๊ฒฐ๊ณผ ๋ฐ ํผ๋๋ฐฑ
"""
base64_image = self.encode_image(image_path)
# ์์ฑ์ด ์์ผ๋ฉด ํ
์คํธ๋ก ๋ณํ
audio_transcript = ""
if audio_path:
audio_transcript = self.transcribe_audio(audio_path)
prompt_text = f"""
๊ณผ๋ชฉ: {subject}
์๋ ์ด๋ฏธ์ง์ ๊ณผ์ ๋ฅผ ํ๊ฐํด์ฃผ์ธ์:
1. ์ ๋ต ์ฌ๋ถ ๋ฐ ์ ์ (100์ ๋ง์ )
2. ์ํ ์
3. ๊ฐ์ ์ด ํ์ํ ์
4. ์์ธํ ํด์ค
5. ์ถ๊ฐ ํ์ต ์๋ฃ ์ถ์ฒ
"""
if audio_transcript:
prompt_text += f"\n\nํ์์ ์ค๋ช
: \"{audio_transcript}\"\n(์ค๋ช
๋ ํจ๊ป ๊ณ ๋ คํด์ฃผ์ธ์)"
messages = [
SystemMessage(content=f"{subject} ๊ต์ก ์ ๋ฌธ๊ฐ AI์
๋๋ค."),
HumanMessage(
content=[
{"type": "text", "text": prompt_text},
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}"
}
}
]
)
]
response = self.llm.invoke(messages)
return response.content
def create_personalized_quiz(self, weak_areas: List[str], difficulty: str):
"""
๋ง์ถคํ ํด์ฆ ์์ฑ
Args:
weak_areas: ์ทจ์ฝํ ์์ญ ๋ฆฌ์คํธ
difficulty: ๋์ด๋ (easy/medium/hard)
Returns:
์์ฑ๋ ํด์ฆ
"""
messages = [
SystemMessage(content="๊ต์ก ์ฝํ
์ธ ์ ์ ์ ๋ฌธ๊ฐ์
๋๋ค."),
HumanMessage(content=f"""
๋ค์ ์์ญ์ ๋ํ {difficulty} ๋์ด๋์ ํด์ฆ๋ฅผ 5๋ฌธ์ ๋ง๋ค์ด์ฃผ์ธ์:
{', '.join(weak_areas)}
๊ฐ ๋ฌธ์ ๋:
1. ๋ฌธ์
2. ์ ํ์ง (4๊ฐ)
3. ์ ๋ต
4. ํด์ค
ํ์์ผ๋ก ์ ๊ณตํด์ฃผ์ธ์.
""")
]
response = self.llm.invoke(messages)
return response.content
# ์ฌ์ฉ ์์
edu_bot = EducationBot()
# ๊ณผ์ ํ๊ฐ
feedback = edu_bot.evaluate_homework(
"math_homework.jpg",
"explanation.mp3",
subject="์ํ"
)
print(f"ํ๊ฐ ๊ฒฐ๊ณผ:\n{feedback}\n")
# ๋ง์ถคํ ํด์ฆ
quiz = edu_bot.create_personalized_quiz(
["์ด์ฐจ๋ฐฉ์ ์", "์ธ์๋ถํด"],
difficulty="medium"
)
print(f"์์ฑ๋ ํด์ฆ:\n{quiz}")
์ฌ๋ฅ๋ท ํ๋ซํผ์์ ์ด๋ฐ ๋ฉํฐ๋ชจ๋ฌ AI๋ฅผ ํ์ฉํ๋ฉด ์ ๋ง ๋ฉ์ง ์๋น์ค๋ฅผ ๋ง๋ค ์ ์์ด:
โข ํฌํธํด๋ฆฌ์ค ์๋ ๋ถ์: ๋์์ด๋๋ ์ํฐ์คํธ์ ์ํ ์ด๋ฏธ์ง๋ฅผ ๋ถ์ํด์ ์คํ์ผ๊ณผ ๊ฐ์ ์ ์๋์ผ๋ก ์์ฝ
โข ์ฌ๋ฅ ๋งค์นญ: ์๋ขฐ์ธ์ด ์ํ๋ ๊ฒฐ๊ณผ๋ฌผ ์ํ๊ณผ ์ค๋ช ์ ์ฌ๋ฆฌ๋ฉด ๊ฐ์ฅ ์ ํฉํ ์ฌ๋ฅ ์ ๊ณต์๋ฅผ ์ถ์ฒ
โข ์์ ๋ฌผ ํ์ง ๊ฒ์: ์์ฑ๋ ์์ ๋ฌผ์ AI๊ฐ ๋จผ์ ๊ฒํ ํด์ ๊ธฐ๋ณธ์ ์ธ ํ์ง ์ฒดํฌ
โข ํ์ต ์ฝํ ์ธ ์ถ์ฒ: ์ฌ์ฉ์์ ๊ด์ฌ์ฌ์ ์์ค์ ๋ถ์ํด์ ๋ง์ถคํ ๊ฐ์ ์ถ์ฒ
โก ์ฑ๋ฅ ์ต์ ํ ๋ฐ ๋น์ฉ ์ ๊ฐ
๋ฉํฐ๋ชจ๋ฌ AI๋ ๊ฐ๋ ฅํ์ง๋ง, ์๋ชป ์ฌ์ฉํ๋ฉด API ๋น์ฉ์ด ์์ฒญ๋๊ฒ ๋์ฌ ์ ์์ด. ํนํ ์ด๋ฏธ์ง์ ์ค๋์ค๋ฅผ ์ฒ๋ฆฌํ ๋๋ ํ ํฐ ์ฌ์ฉ๋์ด ๋ง์์ง๊ฑฐ๋ . ๋๋ํ๊ฒ ์ต์ ํํ๋ ๋ฐฉ๋ฒ์ ์์๋ณด์! ๐ฐ
๐ฏ ์ด๋ฏธ์ง ์ต์ ํ
from PIL import Image
import io
class OptimizedMultimodalBot(AudioMultimodalBot):
def optimize_image(self, image_path: str, max_size=(1024, 1024),
quality=85):
"""
์ด๋ฏธ์ง ์ต์ ํ (ํฌ๊ธฐ ์กฐ์ ๋ฐ ์์ถ)
Args:
image_path: ์๋ณธ ์ด๋ฏธ์ง ๊ฒฝ๋ก
max_size: ์ต๋ ํฌ๊ธฐ (width, height)
quality: JPEG ํ์ง (1-100)
Returns:
์ต์ ํ๋ ์ด๋ฏธ์ง์ base64 ๋ฌธ์์ด
"""
# ์ด๋ฏธ์ง ์ด๊ธฐ
img = Image.open(image_path)
# RGBA๋ฅผ RGB๋ก ๋ณํ (JPEG๋ ํฌ๋ช
๋ ๋ฏธ์ง์)
if img.mode in ('RGBA', 'LA', 'P'):
background = Image.new('RGB', img.size, (255, 255, 255))
if img.mode == 'P':
img = img.convert('RGBA')
background.paste(img, mask=img.split()[-1] if img.mode == 'RGBA' else None)
img = background
# ๋น์จ ์ ์งํ๋ฉฐ ๋ฆฌ์ฌ์ด์ง
img.thumbnail(max_size, Image.Resampling.LANCZOS)
# ๋ฉ๋ชจ๋ฆฌ์ ์ ์ฅ
buffer = io.BytesIO()
img.save(buffer, format='JPEG', quality=quality, optimize=True)
buffer.seek(0)
# base64 ์ธ์ฝ๋ฉ
import base64
return base64.b64encode(buffer.read()).decode('utf-8')
def batch_process_images(self, image_paths: List[str], question: str,
batch_size: int = 3):
"""
์ด๋ฏธ์ง๋ฅผ ๋ฐฐ์น๋ก ๋๋ ์ ์ฒ๋ฆฌ (๋น์ฉ ์ ๊ฐ)
Args:
image_paths: ์ด๋ฏธ์ง ๊ฒฝ๋ก ๋ฆฌ์คํธ
question: ์ง๋ฌธ
batch_size: ํ ๋ฒ์ ์ฒ๋ฆฌํ ์ด๋ฏธ์ง ์
Returns:
๋ฐฐ์น๋ณ ๋ถ์ ๊ฒฐ๊ณผ
"""
results = []
for i in range(0, len(image_paths), batch_size):
batch = image_paths[i:i+batch_size]
# ๋ฐฐ์น ์ฒ๋ฆฌ
encoded_images = [self.optimize_image(img) for img in batch]
content = [
{"type": "text", "text": f"{question} (์ด๋ฏธ์ง {i+1}-{i+len(batch)})"}
]
for encoded_img in encoded_images:
content.append({
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{encoded_img}",
"detail": "low" # ์ ํด์๋ ๋ชจ๋๋ก ๋น์ฉ ์ ๊ฐ
}
})
messages = [
SystemMessage(content="์ด๋ฏธ์ง ๋ฐฐ์น ๋ถ์ ์ ๋ฌธ๊ฐ์
๋๋ค."),
HumanMessage(content=content)
]
response = self.llm.invoke(messages)
results.append({
'batch': i // batch_size + 1,
'images': batch,
'analysis': response.content
})
return results
๐พ ์บ์ฑ ์ ๋ต
import hashlib
import json
from functools import lru_cache
import pickle
class CachedMultimodalBot(OptimizedMultimodalBot):
def __init__(self, model_name="gpt-4-vision-preview", cache_dir="./cache"):
super().__init__(model_name)
self.cache_dir = cache_dir
os.makedirs(cache_dir, exist_ok=True)
def get_cache_key(self, *args):
"""
์บ์ ํค ์์ฑ
"""
# ์ธ์๋ค์ ๋ฌธ์์ด๋ก ๋ณํํ์ฌ ํด์ ์์ฑ
key_string = json.dumps(args, sort_keys=True)
return hashlib.md5(key_string.encode()).hexdigest()
def get_cached_result(self, cache_key: str):
"""
์บ์๋ ๊ฒฐ๊ณผ ๊ฐ์ ธ์ค๊ธฐ
"""
cache_file = os.path.join(self.cache_dir, f"{cache_key}.pkl")
if os.path.exists(cache_file):
with open(cache_file, 'rb') as f:
return pickle.load(f)
return None
def save_to_cache(self, cache_key: str, result):
"""
๊ฒฐ๊ณผ๋ฅผ ์บ์์ ์ ์ฅ
"""
cache_file = os.path.join(self.cache_dir, f"{cache_key}.pkl")
with open(cache_file, 'wb') as f:
pickle.dump(result, f)
def analyze_with_cache(self, image_path: str, question: str):
"""
์บ์ฑ์ ํ์ฉํ ๋ถ์ (๋์ผํ ์์ฒญ์ ์ฌ์ฌ์ฉ)
Args:
image_path: ์ด๋ฏธ์ง ๊ฒฝ๋ก
question: ์ง๋ฌธ
Returns:
๋ถ์ ๊ฒฐ๊ณผ (์บ์ ๋๋ ์๋ก์ด ๋ถ์)
"""
# ์ด๋ฏธ์ง ํด์ ๊ณ์ฐ
with open(image_path, 'rb') as f:
image_hash = hashlib.md5(f.read()).hexdigest()
# ์บ์ ํค ์์ฑ
cache_key = self.get_cache_key(image_hash, question, self.llm.model_name)
# ์บ์ ํ์ธ
cached_result = self.get_cached_result(cache_key)
if cached_result:
print("โ
์บ์์์ ๊ฒฐ๊ณผ๋ฅผ ๊ฐ์ ธ์์ต๋๋ค!")
return cached_result
# ์๋ก์ด ๋ถ์ ์ํ
print("๐ ์๋ก์ด ๋ถ์์ ์ํํฉ๋๋ค...")
result = self.analyze_image_with_text(image_path, question)
# ์บ์์ ์ ์ฅ
self.save_to_cache(cache_key, result)
return result
# ์ฌ์ฉ ์์
cached_bot = CachedMultimodalBot()
# ์ฒซ ๋ฒ์งธ ํธ์ถ - ์ค์ API ํธ์ถ
result1 = cached_bot.analyze_with_cache("product.jpg", "์ด ์ ํ์ ํน์ง์?")
print(result1)
# ๋ ๋ฒ์งธ ํธ์ถ - ์บ์์์ ๊ฐ์ ธ์ด (๋น์ฉ 0)
result2 = cached_bot.analyze_with_cache("product.jpg", "์ด ์ ํ์ ํน์ง์?")
print(result2)
1. ์ด๋ฏธ์ง ํด์๋ ์กฐ์ : "detail": "low" ์ต์ ์ฌ์ฉ ์ ๋น์ฉ์ด ์ฝ 1/3๋ก ๊ฐ์
2. ํ๋กฌํํธ ์ต์ ํ: ๋ถํ์ํ ์ค๋ช ์ ๊ฑฐ, ํต์ฌ๋ง ๊ฐ๊ฒฐํ๊ฒ
3. ๋ฐฐ์น ์ฒ๋ฆฌ: ์ฌ๋ฌ ์์ฒญ์ ํ๋๋ก ๋ฌถ์ด์ ์ฒ๋ฆฌ
4. ์บ์ฑ: ๋์ผํ ์์ฒญ์ ์ฌ์ฌ์ฉ (ํนํ ์์ฃผ ์กฐํ๋๋ ๋ฐ์ดํฐ)
5. ๋ชจ๋ธ ์ ํ: ๊ฐ๋จํ ์์ ์ GPT-3.5 ๊ฐ์ ์ ๋ ดํ ๋ชจ๋ธ ์ฌ์ฉ
6. ์คํธ๋ฆฌ๋ฐ: ๊ธด ์๋ต์ ์คํธ๋ฆฌ๋ฐ์ผ๋ก ๋ฐ์์ ์ฌ์ฉ์ ๊ฒฝํ ๊ฐ์
โ๏ธ ๋น๋๊ธฐ ์ฒ๋ฆฌ๋ก ์๋ ํฅ์
import asyncio
from langchain_openai import ChatOpenAI
class AsyncMultimodalBot(CachedMultimodalBot):
def __init__(self, model_name="gpt-4-vision-preview"):
super().__init__(model_name)
# ๋น๋๊ธฐ LLM ์ด๊ธฐํ
self.async_llm = ChatOpenAI(
model=model_name,
temperature=0.7,
max_tokens=1000
)
async def analyze_async(self, image_path: str, question: str):
"""
๋น๋๊ธฐ ์ด๋ฏธ์ง ๋ถ์
"""
base64_image = self.optimize_image(image_path)
messages = [
SystemMessage(content="์ด๋ฏธ์ง ๋ถ์ ์ ๋ฌธ๊ฐ์
๋๋ค."),
HumanMessage(
content=[
{"type": "text", "text": question},
{
"type": "image_url",
"image_url": {
"url": f"data:image/jpeg;base64,{base64_image}"
}
}
]
)
]
# ๋น๋๊ธฐ ํธ์ถ
response = await self.async_llm.ainvoke(messages)
return response.content
async def analyze_multiple_async(self, tasks: List[tuple]):
"""
์ฌ๋ฌ ์ด๋ฏธ์ง๋ฅผ ๋์์ ๋น๋๊ธฐ ์ฒ๋ฆฌ
Args:
tasks: [(image_path, question), ...] ํํ์ ๋ฆฌ์คํธ
Returns:
๋ชจ๋ ๋ถ์ ๊ฒฐ๊ณผ
"""
# ๋ชจ๋ ์์
์ ๋์์ ์คํ
results = await asyncio.gather(*[
self.analyze_async(img, q) for img, q in tasks
])
return results
# ์ฌ์ฉ ์์
async def main():
bot = AsyncMultimodalBot()
# ์ฌ๋ฌ ์ด๋ฏธ์ง๋ฅผ ๋์์ ์ฒ๋ฆฌ
tasks = [
("image1.jpg", "์ด ์ด๋ฏธ์ง์ ์ฃผ์ ๊ฐ์ฒด๋?"),
("image2.jpg", "์ด ์ด๋ฏธ์ง์ ์์ ํค์?"),
("image3.jpg", "์ด ์ด๋ฏธ์ง์ ๋ถ์๊ธฐ๋?")
]
import time
start = time.time()
results = await bot.analyze_multiple_async(tasks)
end = time.time()
print(f"โฑ๏ธ ์ฒ๋ฆฌ ์๊ฐ: {end - start:.2f}์ด")
for i, result in enumerate(results, 1):
print(f"\n๊ฒฐ๊ณผ {i}:\n{result}")
# ์คํ
# asyncio.run(main())
| ์ฒ๋ฆฌ ๋ฐฉ์ | 3๊ฐ ์ด๋ฏธ์ง ์ฒ๋ฆฌ ์๊ฐ | ์ฅ์ | ๋จ์ |
|---|---|---|---|
| ์์ฐจ ์ฒ๋ฆฌ | ~30์ด | ๊ตฌํ ๊ฐ๋จ, ์์ ์ | ๋๋ฆผ, ๋นํจ์จ์ |
| ๋น๋๊ธฐ ์ฒ๋ฆฌ | ~10์ด | ๋น ๋ฆ, ํจ์จ์ | ๋ณต์ก๋ ์ฆ๊ฐ |
| ๋ฐฐ์น ์ฒ๋ฆฌ | ~12์ด | ๋น์ฉ ์ ๊ฐ | ํ ๋ฒ์ ์ฒ๋ฆฌ ์ ํ |
| ์บ์ฑ ํ์ฉ | ~0.1์ด (์บ์ ํํธ) | ๋งค์ฐ ๋น ๋ฆ, ๋ฌด๋ฃ | ์ ์ฅ ๊ณต๊ฐ ํ์ |
๐ ๋ณด์ ๋ฐ ์๋ฌ ์ฒ๋ฆฌ
์ค์ ์๋น์ค๋ฅผ ๋ง๋ค ๋๋ ๋ณด์๊ณผ ์๋ฌ ์ฒ๋ฆฌ๊ฐ ์ ๋ง ์ค์ํด. ์ฌ์ฉ์ ๋ฐ์ดํฐ๋ฅผ ์์ ํ๊ฒ ๋ณดํธํ๊ณ , ์์์น ๋ชปํ ์ํฉ์๋ ์๋น์ค๊ฐ ์ค๋จ๋์ง ์๋๋ก ํด์ผ ํ๊ฑฐ๋ . ๐ก๏ธ
๐ ๋ณด์ ๊ฐํ
import os
from typing import Optional
import logging
from datetime import datetime, timedelta
import jwt
class SecureMultimodalBot(AsyncMultimodalBot):
def __init__(self, model_name="gpt-4-vision-preview",
max_file_size_mb=10,
allowed_extensions={'.jpg', '.jpeg', '.png', '.mp3', '.wav'}):
super().__init__(model_name)
self.max_file_size = max_file_size_mb * 1024 * 1024 # MB to bytes
self.allowed_extensions = allowed_extensions
# ๋ก๊น
์ค์
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s',
filename='multimodal_bot.log'
)
self.logger = logging.getLogger(__name__)
def validate_file(self, file_path: str) -> tuple[bool, Optional[str]]:
"""
ํ์ผ ์ ํจ์ฑ ๊ฒ์ฌ
Returns:
(์ ํจ ์ฌ๋ถ, ์๋ฌ ๋ฉ์์ง)
"""
# ํ์ผ ์กด์ฌ ํ์ธ
if not os.path.exists(file_path):
return False, "ํ์ผ์ด ์กด์ฌํ์ง ์์ต๋๋ค."
# ํ์ฅ์ ํ์ธ
_, ext = os.path.splitext(file_path)
if ext.lower() not in self.allowed_extensions:
return False, f"ํ์ฉ๋์ง ์๋ ํ์ผ ํ์์
๋๋ค. ํ์ฉ: {self.allowed_extensions}"
# ํ์ผ ํฌ๊ธฐ ํ์ธ
file_size = os.path.getsize(file_path)
if file_size > self.max_file_size:
return False, f"ํ์ผ ํฌ๊ธฐ๊ฐ ๋๋ฌด ํฝ๋๋ค. (์ต๋: {self.max_file_size / 1024 / 1024}MB)"
# ํ์ผ ๋ด์ฉ ๊ฒ์ฆ (์ด๋ฏธ์ง์ธ ๊ฒฝ์ฐ)
if ext.lower() in {'.jpg', '.jpeg', '.png'}:
try:
from PIL import Image
img = Image.open(file_path)
img.verify() # ์์๋ ์ด๋ฏธ์ง ์ฒดํฌ
except Exception as e:
return False, f"์์๋ ์ด๋ฏธ์ง ํ์ผ์
๋๋ค: {str(e)}"
return True, None
def sanitize_input(self, text: str, max_length: int = 5000) -> str:
"""
์
๋ ฅ ํ
์คํธ ์ ์
Args:
text: ์๋ณธ ํ
์คํธ
max_length: ์ต๋ ๊ธธ์ด
Returns:
์ ์ ๋ ํ
์คํธ
"""
# ๊ธธ์ด ์ ํ
text = text[:max_length]
# ์ํํ ๋ฌธ์ ์ ๊ฑฐ (SQL injection, XSS ๋ฑ ๋ฐฉ์ง)
dangerous_chars = ['<script>', '</script>', 'javascript:', 'onerror=']
for char in dangerous_chars:
text = text.replace(char, '')
# ์ฐ์๋ ๊ณต๋ฐฑ ์ ๊ฑฐ
text = ' '.join(text.split())
return text.strip()
def generate_access_token(self, user_id: str, expiry_hours: int = 24) -> str:
"""
์ ๊ทผ ํ ํฐ ์์ฑ (API ์ธ์ฆ์ฉ)
Args:
user_id: ์ฌ์ฉ์ ID
expiry_hours: ๋ง๋ฃ ์๊ฐ (์๊ฐ)
Returns:
JWT ํ ํฐ
"""
secret_key = os.getenv('JWT_SECRET_KEY', 'your-secret-key')
payload = {
'user_id': user_id,
'exp': datetime.utcnow() + timedelta(hours=expiry_hours),
'iat': datetime.utcnow()
}
token = jwt.encode(payload, secret_key, algorithm='HS256')
return token
def verify_access_token(self, token: str) -> tuple[bool, Optional[dict]]:
"""
์ ๊ทผ ํ ํฐ ๊ฒ์ฆ
Returns:
(์ ํจ ์ฌ๋ถ, ํ์ด๋ก๋)
"""
try:
secret_key = os.getenv('JWT_SECRET_KEY', 'your-secret-key')
payload = jwt.decode(token, secret_key, algorithms=['HS256'])
return True, payload
except jwt.ExpiredSignatureError:
return False, None
except jwt.InvalidTokenError:
return False, None
def safe_analyze(self, image_path: str, question: str,
user_token: Optional[str] = None):
"""
์์ ํ ๋ถ์ (๋ชจ๋ ๊ฒ์ฆ ํฌํจ)
Args:
image_path: ์ด๋ฏธ์ง ๊ฒฝ๋ก
question: ์ง๋ฌธ
user_token: ์ฌ์ฉ์ ์ธ์ฆ ํ ํฐ
Returns:
๋ถ์ ๊ฒฐ๊ณผ ๋๋ ์๋ฌ
"""
try:
# ํ ํฐ ๊ฒ์ฆ (์ ํ์ )
if user_token:
is_valid, payload = self.verify_access_token(user_token)
if not is_valid:
self.logger.warning(f"Invalid token attempt")
return {"error": "์ธ์ฆ ์คํจ", "code": 401}
user_id = payload.get('user_id')
else:
user_id = "anonymous"
# ํ์ผ ๊ฒ์ฆ
is_valid, error_msg = self.validate_file(image_path)
if not is_valid:
self.logger.warning(f"File validation failed: {error_msg}")
return {"error": error_msg, "code": 400}
# ์
๋ ฅ ์ ์
clean_question = self.sanitize_input(question)
# ๋ก๊น
self.logger.info(f"User {user_id} analyzing image: {image_path}")
# ๋ถ์ ์ํ
result = self.analyze_image_with_text(image_path, clean_question)
self.logger.info(f"Analysis completed for user {user_id}")
return {
"success": True,
"result": result,
"timestamp": datetime.utcnow().isoformat()
}
except Exception as e:
self.logger.error(f"Error in safe_analyze: {str(e)}", exc_info=True)
return {
"error": "๋ถ์ ์ค ์ค๋ฅ๊ฐ ๋ฐ์ํ์ต๋๋ค.",
"code": 500,
"details": str(e) if os.getenv('DEBUG') == 'True' else None
}
# ์ฌ์ฉ ์์
secure_bot = SecureMultimodalBot()
# ํ ํฐ ์์ฑ
token = secure_bot.generate_access_token("user123")
# ์์ ํ ๋ถ์
result = secure_bot.safe_analyze(
"user_upload.jpg",
"์ด ์ด๋ฏธ์ง๋ฅผ ๋ถ์ํด์ฃผ์ธ์",
user_token=token
)
if result.get("success"):
print(f"๋ถ์ ๊ฒฐ๊ณผ: {result['result']}")
else:
print(f"์๋ฌ: {result['error']}")
๐จ ์๋ฌ ์ฒ๋ฆฌ ๋ฐ ์ฌ์๋ ๋ก์ง
import time
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
from requests.exceptions import RequestException, Timeout
class RobustMultimodalBot(SecureMultimodalBot):
@retry(
stop=stop_after_attempt(3),
wait=wait_exponential(multiplier=1, min=2, max=10),
retry=retry_if_exception_type((RequestException, Timeout))
)
def analyze_with_retry(self, image_path: str, question: str):
"""
์ฌ์๋ ๋ก์ง์ด ํฌํจ๋ ๋ถ์
๋คํธ์ํฌ ์ค๋ฅ๋ ์ผ์์ ์ฅ์ ์ ์๋์ผ๋ก ์ฌ์๋
"""
try:
return self.analyze_image_with_text(image_path, question)
except Exception as e:
self.logger.error(f"Analysis failed: {str(e)}")
raise
def analyze_with_fallback(self, image_path: str, question: str,
fallback_model: str = "gpt-3.5-turbo"):
"""
๋์ฒด ๋ชจ๋ธ์ ์ฌ์ฉํ ํด๋ฐฑ ์ฒ๋ฆฌ
์ฃผ ๋ชจ๋ธ ์คํจ ์ ๋ ์ ๋ ดํ ๋ชจ๋ธ๋ก ์๋ ์ ํ
"""
try:
# ์ฃผ ๋ชจ๋ธ ์๋
return self.analyze_with_retry(image_path, question)
except Exception as e:
self.logger.warning(f"Primary model failed, trying fallback: {str(e)}")
try:
# ํด๋ฐฑ ๋ชจ๋ธ๋ก ์ ํ
original_model = self.llm.model_name
self.llm.model_name = fallback_model
result = self.analyze_image_with_text(image_path, question)
# ์๋ ๋ชจ๋ธ๋ก ๋ณต๊ตฌ
self.llm.model_name = original_model
return {
"result": result,
"fallback_used": True,
"fallback_model": fallback_model
}
except Exception as fallback_error:
self.logger.error(f"Fallback also failed: {str(fallback_error)}")
return {
"error": "๋ชจ๋ ๋ถ์ ์๋๊ฐ ์คํจํ์ต๋๋ค.",
"details": str(fallback_error)
}
def batch_analyze_with_error_handling(self, tasks: List[tuple]):
"""
๋ฐฐ์น ์ฒ๋ฆฌ ์ ๊ฐ๋ณ ์๋ฌ ์ฒ๋ฆฌ
์ผ๋ถ ์์
์ด ์คํจํด๋ ๋๋จธ์ง๋ ๊ณ์ ์งํ
"""
results = []
for idx, (image_path, question) in enumerate(tasks):
try:
result = self.analyze_with_fallback(image_path, question)
results.append({
"index": idx,
"success": True,
"data": result
})
except Exception as e:
self.logger.error(f"Task {idx} failed: {str(e)}")
results.append({
"index": idx,
"success": False,
"error": str(e)
})
# ํต๊ณ ์ ๋ณด ์ถ๊ฐ
success_count = sum(1 for r in results if r["success"])
return {
"results": results,
"statistics": {
"total": len(tasks),
"success": success_count,
"failed": len(tasks) - success_count,
"success_rate": f"{success_count / len(tasks) * 100:.1f}%"
}
}
# ์ฌ์ฉ ์์
robust_bot = RobustMultimodalBot()
# ์ฌ์๋ ๋ก์ง ํ
์คํธ
result = robust_bot.analyze_with_retry("image.jpg", "๋ถ์ํด์ฃผ์ธ์")
# ํด๋ฐฑ ์ฒ๋ฆฌ ํ
์คํธ
result = robust_bot.analyze_with_fallback("image.jpg", "๋ถ์ํด์ฃผ์ธ์")
# ๋ฐฐ์น ์ฒ๋ฆฌ ํ
์คํธ
tasks = [
("image1.jpg", "์ง๋ฌธ1"),
("image2.jpg", "์ง๋ฌธ2"),
("invalid.jpg", "์ง๋ฌธ3"), # ์คํจํ ์์
]
batch_result = robust_bot.batch_analyze_with_error_handling(tasks)
print(f"์ฑ๊ณต๋ฅ : {batch_result['statistics']['success_rate']}")
- โ API ํค๋ฅผ ํ๊ฒฝ๋ณ์๋ก ๊ด๋ฆฌ (.env ํ์ผ ์ฌ์ฉ)
- โ ํ์ผ ์ ๋ก๋ ํฌ๊ธฐ ์ ํ ์ค์
- โ ํ์ฉ๋ ํ์ผ ํ์๋ง ์ฒ๋ฆฌ
- โ ์ฌ์ฉ์ ์ ๋ ฅ ๊ฒ์ฆ ๋ฐ ์ ์
- โ ์๋ฌ ๋ก๊น ๋ฐ ๋ชจ๋ํฐ๋ง
- โ ์ฌ์๋ ๋ก์ง ๊ตฌํ
- โ ํด๋ฐฑ ๋ฉ์ปค๋์ฆ ์ค๋น
- โ ์๋ ์ ํ(Rate Limiting) ์ ์ฉ
- โ HTTPS ์ฌ์ฉ
- โ ์ ๊ธฐ์ ์ธ ๋ณด์ ์ ๋ฐ์ดํธ
๐ ๋ฐฐํฌ ๋ฐ ์๋น์คํ
์ด์ ๋ง๋ ๋ด์ ์ค์ ๋ก ์ฌ์ฉํ ์ ์๋๋ก ๋ฐฐํฌํด๋ณด์! FastAPI๋ฅผ ์ฌ์ฉํด์ REST API๋ก ๋ง๋ค๋ฉด ์น์ด๋ ๋ชจ๋ฐ์ผ ์ฑ์์ ์ฝ๊ฒ ์ฌ์ฉํ ์ ์์ด. ๐
๐ FastAPI๋ก API ์๋ฒ ๋ง๋ค๊ธฐ
๋จผ์ ํ์ํ ํจํค์ง๋ฅผ ์ค์นํ์:
pip install fastapi uvicorn python-multipart
from fastapi import FastAPI, File, UploadFile, Form, HTTPException
from fastapi.middleware.cors import CORSMiddleware
from fastapi.responses import JSONResponse
import shutil
from pathlib import Path
import uuid
app = FastAPI(
title="๋ฉํฐ๋ชจ๋ฌ ๋ถ์ API",
description="์ด๋ฏธ์ง์ ์ค๋์ค๋ฅผ ๋ถ์ํ๋ AI API",
version="1.0.0"
)
# CORS ์ค์ (ํ๋ก ํธ์๋์์ ์ ๊ทผ ๊ฐ๋ฅํ๋๋ก)
app.add_middleware(
CORSMiddleware,
allow_origins=["*"], # ํ๋ก๋์
์์๋ ํน์ ๋๋ฉ์ธ๋ง ํ์ฉ
allow_credentials=True,
allow_methods=["*"],
allow_headers=["*"],
)
# ๋ด ์ธ์คํด์ค ์์ฑ
bot = RobustMultimodalBot()
# ์์ ํ์ผ ์ ์ฅ ๋๋ ํ ๋ฆฌ
UPLOAD_DIR = Path("./uploads")
UPLOAD_DIR.mkdir(exist_ok=True)
@app.get("/")
async def root():
"""
API ์ํ ํ์ธ
"""
return {
"status": "running",
"message": "๋ฉํฐ๋ชจ๋ฌ ๋ถ์ API๊ฐ ์ ์ ์๋ ์ค์
๋๋ค.",
"version": "1.0.0"
}
@app.post("/analyze/image")
async def analyze_image(
file: UploadFile = File(...),
question: str = Form(...),
token: str = Form(None)
):
"""
์ด๋ฏธ์ง ๋ถ์ ์๋ํฌ์ธํธ
Parameters:
- file: ๋ถ์ํ ์ด๋ฏธ์ง ํ์ผ
- question: ์ฌ์ฉ์ ์ง๋ฌธ
- token: ์ธ์ฆ ํ ํฐ (์ ํ)
Returns:
- ๋ถ์ ๊ฒฐ๊ณผ
"""
# ๊ณ ์ ํ์ผ๋ช
์์ฑ
file_id = str(uuid.uuid4())
file_extension = Path(file.filename).suffix
file_path = UPLOAD_DIR / f"{file_id}{file_extension}"
try:
# ํ์ผ ์ ์ฅ
with file_path.open("wb") as buffer:
shutil.copyfileobj(file.file, buffer)
# ๋ถ์ ์ํ
result = bot.safe_analyze(str(file_path), question, token)
# ์์ ํ์ผ ์ญ์
file_path.unlink()
if result.get("error"):
raise HTTPException(
status_code=result.get("code", 500),
detail=result["error"]
)
return JSONResponse(content=result)
except Exception as e:
# ์๋ฌ ๋ฐ์ ์ ํ์ผ ์ ๋ฆฌ
if file_path.exists():
file_path.unlink()
raise HTTPException(status_code=500, detail=str(e))
@app.post("/analyze/audio")
async def analyze_audio(
file: UploadFile = File(...),
analysis_type: str = Form("transcribe") # transcribe, sentiment, summary
):
"""
์ค๋์ค ๋ถ์ ์๋ํฌ์ธํธ
Parameters:
- file: ์ค๋์ค ํ์ผ
- analysis_type: ๋ถ์ ์ ํ (transcribe/sentiment/summary)
Returns:
- ๋ถ์ ๊ฒฐ๊ณผ
"""
file_id = str(uuid.uuid4())
file_extension = Path(file.filename).suffix
file_path = UPLOAD_DIR / f"{file_id}{file_extension}"
try:
with file_path.open("wb") as buffer:
shutil.copyfileobj(file.file, buffer)
if analysis_type == "transcribe":
result = bot.transcribe_audio(str(file_path))
elif analysis_type == "sentiment":
result = bot.analyze_audio_sentiment(str(file_path))
elif analysis_type == "summary":
result = bot.summarize_audio(str(file_path))
else:
raise HTTPException(status_code=400, detail="Invalid analysis type")
file_path.unlink()
return JSONResponse(content={"success": True, "result": result})
except Exception as e:
if file_path.exists():
file_path.unlink()
raise HTTPException(status_code=500, detail=str(e))
@app.post("/analyze/multimodal")
async def analyze_multimodal(
image: UploadFile = File(...),
audio: UploadFile = File(None),
question: str = Form(...)
):
"""
๋ฉํฐ๋ชจ๋ฌ ๋ถ์ ์๋ํฌ์ธํธ (์ด๋ฏธ์ง + ์ค๋์ค)
Parameters:
- image: ์ด๋ฏธ์ง ํ์ผ
- audio: ์ค๋์ค ํ์ผ (์ ํ)
- question: ์ง๋ฌธ
Returns:
- ํตํฉ ๋ถ์ ๊ฒฐ๊ณผ
"""
image_id = str(uuid.uuid4())
image_path = UPLOAD_DIR / f"{image_id}{Path(image.filename).suffix}"
audio_path = None
if audio:
audio_id = str(uuid.uuid4())
audio_path = UPLOAD_DIR / f"{audio_id}{Path(audio.filename).suffix}"
try:
# ์ด๋ฏธ์ง ์ ์ฅ
with image_path.open("wb") as buffer:
shutil.copyfileobj(image.file, buffer)
# ์ค๋์ค ์ ์ฅ (์๋ ๊ฒฝ์ฐ)
if audio and audio_path:
with audio_path.open("wb") as buffer:
shutil.copyfileobj(audio.file, buffer)
result = bot.multimodal_analysis(
str(image_path),
str(audio_path),
question
)
else:
result = bot.analyze_image_with_text(str(image_path), question)
# ํ์ผ ์ ๋ฆฌ
image_path.unlink()
if audio_path and audio_path.exists():
audio_path.unlink()
return JSONResponse(content={"success": True, "result": result})
except Exception as e:
# ์๋ฌ ์ ํ์ผ ์ ๋ฆฌ
if image_path.exists():
image_path.unlink()
if audio_path and audio_path.exists():
audio_path.unlink()
raise HTTPException(status_code=500, detail=str(e))
@app.post("/auth/token")
async def create_token(user_id: str = Form(...)):
"""
์ธ์ฆ ํ ํฐ ์์ฑ
Parameters:
- user_id: ์ฌ์ฉ์ ID
Returns:
- JWT ํ ํฐ
"""
token = bot.generate_access_token(user_id)
return {"token": token, "expires_in": "24h"}
if __name__ == "__main__":
import uvicorn
uvicorn.run(app, host="0.0.0.0", port=8000)
์๋ฒ๋ฅผ ์คํํ๋ ค๋ฉด:
python api_server.py
# ๋๋
uvicorn api_server:app --reload --host 0.0.0.0 --port 8000
์ด์ ๋ธ๋ผ์ฐ์ ์์ http://localhost:8000/docs๋ก ์ ์ํ๋ฉด ์๋์ผ๋ก ์์ฑ๋ API ๋ฌธ์๋ฅผ ๋ณผ ์ ์์ด! FastAPI์ Swagger UI๊ฐ ์ ๋ง ํธ๋ฆฌํ์ง? ๐
๐ณ Docker๋ก ์ปจํ ์ด๋ํ
๋ฐฐํฌ๋ฅผ ์ฝ๊ฒ ํ๋ ค๋ฉด Docker๋ฅผ ์ฌ์ฉํ๋ ๊ฒ ์ข์. Dockerfile์ ๋ง๋ค์ด๋ณด์:
# Dockerfile
FROM python:3.10-slim
WORKDIR /app
# ์์คํ
ํจํค์ง ์ค์น
RUN apt-get update && apt-get install -y \
ffmpeg \
libsm6 \
libxext6 \
&& rm -rf /var/lib/apt/lists/*
# Python ํจํค์ง ์ค์น
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# ์ ํ๋ฆฌ์ผ์ด์
์ฝ๋ ๋ณต์ฌ
COPY . .
# ํฌํธ ๋
ธ์ถ
EXPOSE 8000
# ์๋ฒ ์คํ
CMD ["uvicorn", "api_server:app", "--host", "0.0.0.0", "--port", "8000"]
requirements.txt ํ์ผ๋ ๋ง๋ค์ด์ผ ํด:
langchain==0.1.0
langchain-openai==0.0.5
langchain-anthropic==0.1.0
fastapi==0.109.0
uvicorn==0.27.0
python-multipart==0.0.6
pillow==10.2.0
python-dotenv==1.0.0
openai-whisper==20231117
pydub==0.25.1
tenacity==8.2.3
pyjwt==2.8.0
Docker ์ด๋ฏธ์ง ๋น๋ ๋ฐ ์คํ:
# ์ด๋ฏธ์ง ๋น๋
docker build -t multimodal-bot:latest .
# ์ปจํ
์ด๋ ์คํ
docker run -d \
--name multimodal-bot \
-p 8000:8000 \
-e OPENAI_API_KEY=your_key_here \
-v $(pwd)/uploads:/app/uploads \
multimodal-bot:latest
# ๋ก๊ทธ ํ์ธ
docker logs -f multimodal-bot
โ๏ธ ํด๋ผ์ฐ๋ ๋ฐฐํฌ (AWS ์์)
1๋จ๊ณ: EC2 ์ธ์คํด์ค ์์ฑ (Ubuntu 22.04, t3.medium ์ด์ ๊ถ์ฅ)
2๋จ๊ณ: ๋ณด์ ๊ทธ๋ฃน์์ 8000๋ฒ ํฌํธ ๊ฐ๋ฐฉ
3๋จ๊ณ: SSH๋ก ์ ์ ํ Docker ์ค์น
4๋จ๊ณ: ์ฝ๋ ์ ๋ก๋ ๋ฐ Docker ์ด๋ฏธ์ง ๋น๋
5๋จ๊ณ: ์ปจํ ์ด๋ ์คํ
6๋จ๊ณ: Nginx๋ก ๋ฆฌ๋ฒ์ค ํ๋ก์ ์ค์ (HTTPS ์ ์ฉ)
# Nginx ์ค์ ์์ (/etc/nginx/sites-available/multimodal-bot)
server {
listen 80;
server_name your-domain.com;
location / {
proxy_pass http://localhost:8000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
# ํ์ผ ์
๋ก๋ ํฌ๊ธฐ ์ ํ
client_max_body_size 50M;
}
}
# HTTPS ์ค์ (Let's Encrypt)
# sudo certbot --nginx -d your-domain.com
๐ ๋ชจ๋ํฐ๋ง ๋ฐ ๋ถ์
์๋น์ค๋ฅผ ์ด์ํ๋ค ๋ณด๋ฉด ์ฑ๋ฅ ๋ชจ๋ํฐ๋ง๊ณผ ์ฌ์ฉ ํจํด ๋ถ์์ด ํ์ํด. ๊ฐ๋จํ ๋ชจ๋ํฐ๋ง ์์คํ ์ ์ถ๊ฐํด๋ณด์! ๐
from datetime import datetime
import json
from collections import defaultdict
from fastapi import Request
import time
class AnalyticsMiddleware:
def __init__(self):
self.stats = defaultdict(lambda: {
'count': 0,
'total_time': 0,
'errors': 0
})
async def __call__(self, request: Request, call_next):
start_time = time.time()
try:
response = await call_next(request)
# ํต๊ณ ์
๋ฐ์ดํธ
endpoint = request.url.path
process_time = time.time() - start_time
self.stats[endpoint]['count'] += 1
self.stats[endpoint]['total_time'] += process_time
# ์๋ต ํค๋์ ์ฒ๋ฆฌ ์๊ฐ ์ถ๊ฐ
response.headers["X-Process-Time"] = str(process_time)
return response
except Exception as e:
endpoint = request.url.path
self.stats[endpoint]['errors'] += 1
raise
def get_stats(self):
"""
ํต๊ณ ์ ๋ณด ๋ฐํ
"""
result = {}
for endpoint, data in self.stats.items():
avg_time = data['total_time'] / data['count'] if data['count'] > 0 else 0
result[endpoint] = {
'requests': data['count'],
'avg_response_time': f"{avg_time:.3f}s",
'errors': data['errors'],
'error_rate': f"{data['errors'] / data['count'] * 100:.1f}%" if data['count'] > 0 else "0%"
}
return result
# FastAPI ์ฑ์ ๋ฏธ๋ค์จ์ด ์ถ๊ฐ
analytics = AnalyticsMiddleware()
app.middleware("http")(analytics)
@app.get("/stats")
async def get_statistics():
"""
API ์ฌ์ฉ ํต๊ณ ์กฐํ
"""
return analytics.get_stats()
๐ ๋ง๋ฌด๋ฆฌํ๋ฉฐ
์! ์ฌ๊ธฐ๊น์ง ์ ๋ง ๊ธด ์ฌ์ ์ด์์ด. ์ฐ๋ฆฌ๋ LangChain์ ์ฌ์ฉํด์ ํ
์คํธ, ์ด๋ฏธ์ง, ์ค๋์ค๋ฅผ ๋ชจ๋ ์ฒ๋ฆฌํ ์ ์๋ ๋ฉํฐ๋ชจ๋ฌ AI ๋ด์ ๋ง๋ค์์ด. ๊ธฐ๋ณธ์ ์ธ ๊ตฌํ๋ถํฐ ์์ํด์ ๋ณด์, ์ต์ ํ, ๋ฐฐํฌ๊น์ง ๋ชจ๋ ๊ณผ์ ์ ๋ค๋ค์ง. ๐
์ด์ ๋๋ง์ ๋ฉํฐ๋ชจ๋ฌ ๋ด์ ๋ง๋ค ์ค๋น๊ฐ ๋์ด! ์๋ฃ, ์ด์ปค๋จธ์ค, ๊ต์ก ๋ฑ ๋ค์ํ ๋ถ์ผ์ ์ ์ฉํ ์ ์๊ณ , ์ฌ๋ฅ๋ท ๊ฐ์ ํ๋ซํผ์์๋ ํ์ฉํ ์ ์๊ฒ ์ง? ์ฌ์ฉ์๋ค์๊ฒ ๋ ๋์ ๊ฒฝํ์ ์ ๊ณตํ๋ ๋๋ํ AI ์๋น์ค๋ฅผ ๋ง๋ค์ด๋ณด์! ๐ช
1. ์คํํด๋ณด๊ธฐ: ๋ค์ํ ํ๋กฌํํธ์ ํ๋ผ๋ฏธํฐ๋ฅผ ํ ์คํธํ๋ฉฐ ์ต์ ์ ์ค์ ์ฐพ๊ธฐ
2. ํ์ฅํ๊ธฐ: ๋น๋์ค ๋ถ์, ์ค์๊ฐ ์คํธ๋ฆฌ๋ฐ ๋ฑ ์๋ก์ด ๊ธฐ๋ฅ ์ถ๊ฐ
3. ์ต์ ํํ๊ธฐ: ์บ์ฑ, ๋ฐฐ์น ์ฒ๋ฆฌ ๋ฑ์ผ๋ก ๋น์ฉ๊ณผ ์๋ ๊ฐ์
4. ์ปค๋ฎค๋ํฐ ์ฐธ์ฌ: LangChain GitHub, Discord์์ ๋ค๋ฅธ ๊ฐ๋ฐ์๋ค๊ณผ ๊ต๋ฅ
5. ํ๋ก์ ํธ ๊ณต์ : ์ฌ๋ฅ๋ท์์ ๋์ AI ๊ฐ๋ฐ ์ฌ๋ฅ์ ๊ณต์ ํ๊ณ ์์ตํํ๊ธฐ! ๐ฐ
๋ฉํฐ๋ชจ๋ฌ AI๋ ๊ณ์ ๋ฐ์ ํ๊ณ ์์ด. GPT-5, Gemini 2.0 ๊ฐ์ ์๋ก์ด ๋ชจ๋ธ๋ค์ด ๋์ค๋ฉด์ ๋ ๊ฐ๋ ฅํ ๊ธฐ๋ฅ๋ค์ด ์ถ๊ฐ๋๊ณ ์๊ฑฐ๋ . ์ง๊ธ ๋ฐฐ์ด ๊ธฐ์ด๋ฅผ ๋ฐํ์ผ๋ก ๊ณ์ ํ์ตํ๊ณ ์คํํ๋ฉด, ์ ๋ง ๋ฉ์ง AI ์๋น์ค๋ฅผ ๋ง๋ค ์ ์์ ๊ฑฐ์ผ! ๐
๊ถ๊ธํ ์ ์ด ์๊ฑฐ๋ ๋งํ๋ ๋ถ๋ถ์ด ์์ผ๋ฉด LangChain ๊ณต์ ๋ฌธ์๋ ์ปค๋ฎค๋ํฐ๋ฅผ ํ์ฉํด๋ด. ๊ทธ๋ฆฌ๊ณ ์ฌ๋ฅ๋ท์์ AI ๊ฐ๋ฐ ๊ด๋ จ ์ฌ๋ฅ์ ์ฐพ์๋ณด๋ ๊ฒ๋ ์ข์ ๋ฐฉ๋ฒ์ด์ผ. ํจ๊ป ๋ฐฐ์ฐ๊ณ ์ฑ์ฅํ๋ ๊ฒ ๊ฐ์ฅ ๋น ๋ฅธ ๊ธธ์ด๋๊น! ๐
์, ์ด์ ์ฝ๋ฉ์ ์์ํด๋ณผ๊น? ํ์ดํ
! ๐ฅ
๐ ์ถํํด์! ๋ฉํฐ๋ชจ๋ฌ AI ๋ด ๋ง์คํฐ๊ฐ ๋์์ด์!
์ด์ ๋น์ ๋ง์ ๋๋ํ AI ์๋น์ค๋ฅผ ๋ง๋ค์ด๋ณด์ธ์.
์ฌ๋ฅ๋ท์์ ์ฌ๋ฌ๋ถ์ AI ๊ฐ๋ฐ ์ฌ๋ฅ์ ๊ณต์ ํ๊ณ ์ฑ์ฅํ์ธ์! ๐
๊ด๋ จ ํค์๋
๋๊ธ 0
์ง์์ธ์ ์ฒ - ์ง์ ์ฌ์ฐ๊ถ ๋ณดํธ ๊ณ ์ง
์ง์ ์ฌ์ฐ๊ถ ๋ณดํธ ๊ณ ์ง
- ์ ์๊ถ ๋ฐ ์์ ๊ถ: ๋ณธ ์ปจํ ์ธ ๋ ์ฌ๋ฅ๋ท์ ๋ ์ AI ๊ธฐ์ ๋ก ์์ฑ๋์์ผ๋ฉฐ, ๋ํ๋ฏผ๊ตญ ์ ์๊ถ๋ฒ ๋ฐ ๊ตญ์ ์ ์๊ถ ํ์ฝ์ ์ํด ๋ณดํธ๋ฉ๋๋ค.
- AI ์์ฑ ์ปจํ ์ธ ์ ๋ฒ์ ์ง์: ๋ณธ AI ์์ฑ ์ปจํ ์ธ ๋ ์ฌ๋ฅ๋ท์ ์ง์ ์ฐฝ์๋ฌผ๋ก ์ธ์ ๋๋ฉฐ, ๊ด๋ จ ๋ฒ๊ท์ ๋ฐ๋ผ ์ ์๊ถ ๋ณดํธ๋ฅผ ๋ฐ์ต๋๋ค.
- ์ฌ์ฉ ์ ํ: ์ฌ๋ฅ๋ท์ ๋ช ์์ ์๋ฉด ๋์ ์์ด ๋ณธ ์ปจํ ์ธ ๋ฅผ ๋ณต์ , ์์ , ๋ฐฐํฌ, ๋๋ ์์ ์ ์ผ๋ก ํ์ฉํ๋ ํ์๋ ์๊ฒฉํ ๊ธ์ง๋ฉ๋๋ค.
- ๋ฐ์ดํฐ ์์ง ๊ธ์ง: ๋ณธ ์ปจํ ์ธ ์ ๋ํ ๋ฌด๋จ ์คํฌ๋ํ, ํฌ๋กค๋ง, ๋ฐ ์๋ํ๋ ๋ฐ์ดํฐ ์์ง์ ๋ฒ์ ์ ์ฌ์ ๋์์ด ๋ฉ๋๋ค.
- AI ํ์ต ์ ํ: ์ฌ๋ฅ๋ท์ AI ์์ฑ ์ปจํ ์ธ ๋ฅผ ํ AI ๋ชจ๋ธ ํ์ต์ ๋ฌด๋จ ์ฌ์ฉํ๋ ํ์๋ ๊ธ์ง๋๋ฉฐ, ์ด๋ ์ง์ ์ฌ์ฐ๊ถ ์นจํด๋ก ๊ฐ์ฃผ๋ฉ๋๋ค.

๋๊ธ ์์ฑ
์ด ๊ธ์ ๋ํ ์ฌ๋ฌ๋ถ์ ์๊ฐ์ ๋ค๋ ค์ฃผ์ธ์
๋ก๊ทธ์ธ์ด ํ์ํฉ๋๋ค
๋๊ธ์ ์์ฑํ๋ ค๋ฉด ๋จผ์ ๋ก๊ทธ์ธํด์ฃผ์ธ์.