What Is Nividia's Chat With RTX, And How To Use It

Опубликовано: 05 Октябрь 2024
на канале: Techmeme Ride Home Podcast
23,269
218

Nvidia has released an early version of Chat with RTX, an app that lets users run a personal AI chatbot on a PC with an RTX 30- or 40-series GPU and 8GB+ of VRAM. Quoting Venture Beat:
The new offering, Chat with RTX, allows users to harness the power of personalized generative AI directly on their local devices, showcasing the potential of retrieval-augmented generation (RAG) and TensorRT-LLM software. At the same time, it doesn’t burn up a lot of data center computing and it helps with local privacy so that users don’t have to worry about their AI chats.

Nvidia said that Chat with RTX is more than a mere chatbot; it’s a personalized AI companion that users can customize with their own content. By leveraging the capabilities of local GeForce-powered Windows PCs, users can accelerate their experience and enjoy the benefits of generative AI with unprecedented speed and privacy, the company said.

The tool leverages RAG, TensorRT-LLM software, and Nvidia RTX acceleration to facilitate quick, contextually relevant answers based on local datasets. Users can connect the application to local files on their PCs, turning them into a dataset for open-source large language models like Mistral or Llama 2.

Rather than sifting through various files, users can type natural language queries, such as asking about a restaurant recommendation or any personalized information, and Chat with RTX will swiftly scan and provide the answer with context. The application supports a variety of file formats, including .txt, .pdf, .doc/.docx, and .xml, making it versatile and user-friendly. ENDQUOTE

Quoting HowToGeek:

Because Chat With RTX runs locally, it produces fast results without sending your personal data into the cloud. The LLM will only scan files or folders that are selected by the user. I should note that other LLMs, including those from HuggingFace and OpenAI, can run locally. Chat With RTX is notable for two reasons—it doesn't require any expertise, and it shows the capabilities of NVIDIA's open-source TensorRT-LLM RAG, which developers can use to build their own AI applications. ENDQUOTE

And quoting The Verge:

I’ve been briefly testing out Chat with RTX over the past day, and although the app is a little rough around the edges, I can already see this being a valuable part of data research for journalists or anyone who needs to analyze a collection of documents.

Chat with RTX can handle YouTube videos, so you simply input a URL, and it lets you search transcripts for specific mentions or summarize an entire video. I found this ideal for searching through video podcasts, particularly for finding specific mentions in podcasts over the past week amid rumors of Microsoft’s new Xbox strategy shift.

When it worked properly I was able to find references in videos within seconds. I also created a dataset of FTC v. Microsoft documents for Chat with RTX to analyze. When I was covering the court case last year, it was often overwhelming to search through documents at speed, but Chat with RTX helped me query them nearly instantly on my PC.

I’ve also found this useful to scan through PDFs and fact-check data. Microsoft’s own Copilot system doesn’t handle PDFs well within Word, but Nvidia’s Chat with RTX had no problem pulling out all the key information. The responses are near instant as well, with none of the lag you usually see when using cloud-based ChatGPT or Copilot chatbots. ENDDQUOTE

All the early reviews of using this stress this is probably more a demo than a fully fledged product. It takes around 30 minutes to install, you have to then have Mistral or Llama 2 installed to query the data. And there are tons of reports of bugs and problems getting it to even run. But if you want to experiment with the cutting edge of edge AI, knock yourself out.

#ainews #nvidia #technews