Загрузка видео...

Не удалось загрузить видео

На главную

Fully local Code Assistant running on NVIDIA GPU! In this tutorial, I'll show you how to run Llama3 using TensorRT and Nvidia's Triton Inference Server to use it as a Code Assistant in VSCode In this thread 🧵, I'll walk you through the integration process, explaining each step simply...

42,154 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 11

Фото профиля Daniel San
Daniel San2 лет назад

To get started, we need a @nvidia GPU 🤩 In this case, we will use the following hardware 💻

Фото профиля Daniel San
Daniel San2 лет назад

We need to have Docker and CUDA installed Follow the guides below for installing both tools Docker: CUDA: and then run the following commands to confirm everything is set up correctly

Фото профиля Daniel San
Daniel San2 лет назад

Download the llama3-8B model from @huggingface

Фото профиля Daniel San
Daniel San2 лет назад

Now, Run TensorRT to compile the model using the Docker container Clone the TensorRT repository and move the model folder

Фото профиля Daniel San
Daniel San2 лет назад

You should now be able to test the compiled model

Фото профиля Daniel San
Daniel San2 лет назад

Perfect! We have the model now, let's deploy it on Triton Inference Server

Фото профиля Daniel San
Daniel San2 лет назад

The server is up and ready to connect with CodeGPT via the custom connection Open CodeGPT in VSCode, select Custom as the provider, and enter "ensemble" for the model

Фото профиля Daniel San
Daniel San2 лет назад

That's all! I'm sharing the link to the full article with all the details of the tutorial

Фото профиля Alexander Mia
Alexander Mia1 год назад

INTRODUCING: Agentic Security - LLM Security Scanner! 🔍 🔑 Features: Scans for prompt injections, jailbreaking & more. Provides detailed reports & options to customize attack rules. 🔗access the GitHub Link ↓

Фото профиля ₣rancisco Trillo
₣rancisco Trillo2 лет назад

Or just use Continue and Ollama with whatever brand GPU 🤷‍♂️ that’s open source

Фото профиля Daniel San
Daniel San2 лет назад

you can also use CodeGPT with Ollama Check this link:

Похожие видео

AMD might have disrupted Nvidia's entire cloud GPU rental business. In January at CES, AMD CEO Lisa Su demonstrated a $1,499 mini PC running the same class of AI model that currently costs companies $2,500 to $3,000 every month to rent from Nvidia-powered cloud servers. AMD's own branded version opened pre-orders this month at $3,999. Third party manufacturers have been selling the same chip since 2025 starting at $1,499. Here is exactly why this is dangerous for Nvidia. Nvidia's $75 billion quarterly revenue is built almost entirely on one business model, companies rent access to Nvidia GPUs through cloud providers like AWS and Lambda Labs to run AI. They pay monthly. Nvidia gets paid every time someone runs an AI model in the cloud. That recurring rental income is what turned Nvidia into a $5 trillion company. The AMD box eliminates that monthly fee permanently. One AI consultant switched from $2,800 per month in Nvidia cloud rental costs to $8 per month in electricity. The hardware paid for itself in 11 days. Over 8 months he generated $47,000 running the same AI workloads that previously left him paying Nvidia's ecosystem $2,800 every single month. Multiply that across thousands of enterprise customers and the revenue erosion becomes structural. Every business that buys this box stops paying cloud rental fees forever. Lawyers, doctors, banks, accountants, and financial advisors, businesses with sensitive data that cannot legally go to a cloud server represent billions in annual cloud GPU fees that Nvidia is now at risk of losing permanently. The threat is also closing in from the top. Google signed deals worth tens of billions with Anthropic and Meta to replace Nvidia with its own chips. Amazon built its own AI chips across AWS. Apple trained its AI on Google's chips, not Nvidia's. Custom silicon has grown from 21% of the AI chip market in 2025 to 28% in 2026. Nvidia's rental model only worked because serious AI compute had no alternative.

Bull Theory

26,668 просмотров • 1 месяц назад