Ser*_*sov 5 python pyngrok ollama
我有一个这样的代码。我正在启动它。我得到一个 ngrok 链接。
!pip install aiohttp pyngrok
import os
import asyncio
from aiohttp import ClientSession
# Set LD_LIBRARY_PATH so the system NVIDIA library becomes preferred
# over the built-in library. This is particularly important for
# Google Colab which installs older drivers
os.environ.update({'LD_LIBRARY_PATH': '/usr/lib64-nvidia'})
async def run(cmd):
'''
run is a helper function to run subcommands asynchronously.
'''
print('>>> starting', *cmd)
p = await asyncio.subprocess.create_subprocess_exec(
*cmd,
stdout=asyncio.subprocess.PIPE,
stderr=asyncio.subprocess.PIPE,
)
async def pipe(lines):
async for line in lines:
print(line.strip().decode('utf-8'))
await asyncio.gather(
pipe(p.stdout),
pipe(p.stderr),
)
await asyncio.gather(
run(['ollama', 'serve']),
run(['ngrok', 'http', '--log', 'stderr', '11434']),
)
Run Code Online (Sandbox Code Playgroud)
我正在关注,但以下内容在页面上
我怎样才能解决这个问题?在此之前,我做了以下事情
!choco install ngrok
!ngrok config add-authtoken -----
Run Code Online (Sandbox Code Playgroud)
!curl https://ollama.ai/install.sh | sh
!command -v systemctl >/dev/null && sudo systemctl stop ollama
Run Code Online (Sandbox Code Playgroud)
!curl https://ollama.ai/install.sh | sh
# should produce, among other thigns:
# The Ollama API is now available at 0.0.0.0:11434
Run Code Online (Sandbox Code Playgroud)
这意味着 Ollama 正在运行(但请检查是否存在错误,特别是在图形功能/Cuda 方面,因为这些可能会产生干扰。
但是,不要跑
!command -v systemctl >/dev/null && sudo systemctl stop ollama
(除非你想阻止奥拉玛)。
下一步是启动 Ollama 服务,但既然您正在使用,ngrok我假设您希望能够从 Colab 之外的其他环境运行 LLM?如果情况并非如此,那么您实际上并不需要 ngrok,但由于 Colab 很难与异步代码和线程很好地工作,因此使用 Colab 来运行一个足够强大的 VM 来处理比(比如说)您可以在开发环境中运行的东西(如果这是一个问题)。
Ollama 尚未作为服务运行,但我们可以提前设置 ngrok:
!curl https://ollama.ai/install.sh | sh
# should produce, among other thigns:
# The Ollama API is now available at 0.0.0.0:11434
Run Code Online (Sandbox Code Playgroud)
运行该代码,以便函数存在,然后在下一个单元格中,在单独的线程中启动 ngrok,这样它就不会挂起您的 colab - 我们将使用队列,以便我们仍然可以在线程之间共享数据,因为我们想知道ngrok 运行时的公共 URL 将是:
import threading
import time
import os
import asyncio
from pyngrok import ngrok
import threading
import queue
import time
from threading import Thread
# Get your ngrok token from your ngrok account:
# https://dashboard.ngrok.com/get-started/your-authtoken
token="your token goes here - don't forget to replace this with it!"
ngrok.set_auth_token(token)
# set up a stoppable thread (not mandatory, but cleaner if you want to stop this later
class StoppableThread(threading.Thread):
def __init__(self, *args, **kwargs):
super(StoppableThread, self).__init__(*args, **kwargs)
self._stop_event = threading.Event()
def stop(self):
self._stop_event.set()
def is_stopped(self):
return self._stop_event.is_set()
def start_ngrok(q, stop_event):
try:
# Start an HTTP tunnel on the specified port
public_url = ngrok.connect(11434)
# Put the public URL in the queue
q.put(public_url)
# Keep the thread alive until stop event is set
while not stop_event.is_set():
time.sleep(1) # Adjust sleep time as needed
except Exception as e:
print(f"Error in start_ngrok: {e}")
Run Code Online (Sandbox Code Playgroud)
它将运行,但您需要从队列中获取结果以查看 ngrok 返回的内容,因此请执行以下操作:
# Create a queue to share data between threads
url_queue = queue.Queue()
# Start ngrok in a separate thread
ngrok_thread = StoppableThread(target=start_ngrok, args=(url_queue, StoppableThread.is_stopped))
ngrok_thread.start()
Run Code Online (Sandbox Code Playgroud)
这应该输出类似:
Ngrok tunnel established at: NgrokTunnel: "https://{somelongsubdomain}.ngrok-free.app" -> "http://localhost:11434"
Run Code Online (Sandbox Code Playgroud)
# Wait for the ngrok tunnel to be established
while True:
try:
public_url = url_queue.get()
if public_url:
break
print("Waiting for ngrok URL...")
time.sleep(1)
except Exception as e:
print(f"Error in retrieving ngrok URL: {e}")
print("Ngrok tunnel established at:", public_url)
Run Code Online (Sandbox Code Playgroud)
这将创建运行异步命令的函数,但尚未运行它。
这将在一个单独的线程中启动 ollama,这样你的 Colab 就不会被阻塞:
Ngrok tunnel established at: NgrokTunnel: "https://{somelongsubdomain}.ngrok-free.app" -> "http://localhost:11434"
Run Code Online (Sandbox Code Playgroud)
它应该产生类似的结果:
import os
import asyncio
# NB: You may need to set these depending and get cuda working depending which backend you are running.
# Set environment variable for NVIDIA library
# Set environment variables for CUDA
os.environ['PATH'] += ':/usr/local/cuda/bin'
# Set LD_LIBRARY_PATH to include both /usr/lib64-nvidia and CUDA lib directories
os.environ['LD_LIBRARY_PATH'] = '/usr/lib64-nvidia:/usr/local/cuda/lib64'
async def run_process(cmd):
print('>>> starting', *cmd)
process = await asyncio.create_subprocess_exec(
*cmd,
stdout=asyncio.subprocess.PIPE,
stderr=asyncio.subprocess.PIPE
)
# define an async pipe function
async def pipe(lines):
async for line in lines:
print(line.decode().strip())
await asyncio.gather(
pipe(process.stdout),
pipe(process.stderr),
)
# call it
await asyncio.gather(pipe(process.stdout), pipe(process.stderr))
Run Code Online (Sandbox Code Playgroud)
现在你已经准备好了。您可以在 Colab 中执行后续步骤,但如果您通常在本地计算机上进行开发,那么在本地计算机上运行可能会更容易。
假设您已经在本地开发环境(例如 WSL2)上安装了 ollama,我假设它是 linux……但即您面前的笔记本电脑或台式机(而不是 Colab)。
将下面的实际 URI 替换为上面 ngrok 报告的任何公共 URI:
import asyncio
import threading
async def start_ollama_serve():
await run_process(['ollama', 'serve'])
def run_async_in_thread(loop, coro):
asyncio.set_event_loop(loop)
loop.run_until_complete(coro)
loop.close()
# Create a new event loop that will run in a new thread
new_loop = asyncio.new_event_loop()
# Start ollama serve in a separate thread so the cell won't block execution
thread = threading.Thread(target=run_async_in_thread, args=(new_loop, start_ollama_serve()))
thread.start()
Run Code Online (Sandbox Code Playgroud)
您现在可以运行 ollama,它将在您的 Colab 中的遥控器上运行(只要它保持正常运行)。
例如,在本地计算机上运行它,它看起来好像在本地运行,但它实际上在您的 Colab 中运行,并且结果将提供给您调用它的任何地方(只要 OLLAMA_HOST 设置正确并且是到的有效隧道)您的乌拉马服务:
>>> starting ollama serve
Couldn't find '/root/.ollama/id_ed25519'. Generating new private key.
Your new public key is:
ssh-ed25519 {some key}
2024/01/16 20:19:11 images.go:808: total blobs: 0
2024/01/16 20:19:11 images.go:815: total unused blobs removed: 0
2024/01/16 20:19:11 routes.go:930: Listening on 127.0.0.1:11434 (version 0.1.20)
Run Code Online (Sandbox Code Playgroud)
您现在可以在本地命令行上与模型交互,但模型在 Colab 上运行。
如果您想运行更大的模型,例如 mixtral,那么您需要确保将 Colab 连接到足够强大的后端计算(例如 48GB+ 的 RAM,因此在撰写本文时 V100 GPU 是此功能的最低规格)。
注意:如果上述任何步骤的输出中显示 cuda 或 nvidia 存在任何问题,请在解决这些问题之前不要继续。
希望有帮助!
粗暴