Introducing AnythingLLM with NPU Support
Hi there. My name is Timothy Kabat, founder of Mintplex Labs and creator of AnythingLLM.
AnythingLLM is an all-in-one desktop application to leverage AI to chat with documents, run AI agents, and do many, many, many more things, all locally and privately on device. And I'm happy to announce that our latest build of AnythingLLM comes pre-built with support for running LLMs and other models that we use directly on the NPU for the fastest performance and power efficiency that you can get on the market today. This video is going to be a brief overview of what AnythingLLM can do and also a short demo of how powerful it really is and can be for you in your use case.
Customization and Multi-Provider LLM Support
AnythingLLM likes to try to be as customizable and built to your use case as it possibly can be right out of the box. And for that, we support many different languages and other different styles and themes that you can use that just fit maybe your use case or your branding or just what you like.
One of the most important features, and the newest, is support for running LLMs directly optimized for the NPU on Snapdragon X and Copilot PC devices. However, we're not limited to just a couple models. Some of your favorite providers that you already know, if you want to run any model that's out there on the web on CPU, we can support that, but also running and connecting with popular cloud models like OpenAI, Anthropic, Gemini, also other local LLM providers you may have a preference for like LM Studio. All of these are possible right here, built in, and you can actually use more than one provider at the same time and across very different workspaces so that you get the best experience no matter what.
Privacy, On-Device Storage, and NPU-Optimized Embeddings
Of course, everything by default is private and on device. And for that, we also have a built-in vector database so that your storage of your proprietary information or the documents that you're chatting with stay on your device. And for that, we have LanceDB. It works right out of the box. And I'm also happy to announce that for embedding those documents that you will be using with your LLM, we also have optimized our embedding model for the NPU as well. So you'll get the fastest speed that you can get on the market all across the board when you use AnythingLLM on a Copilot PC.
Demo: Creating a Workspace & Chatting on the NPU
One of the most common use cases of AnythingLLM is chatting with a document. So to do that, we're going to create what is called a workspace. A workspace is just a collection of documents, resources, tools, and just any other things that an LLM can use to help answer your question. So let's create a workspace.
Now, as you can see, workspaces have a very familiar UI. And for this, we're just going to want to chat with our LLM directly, no uploaded context or anything fancy. Let's just have a conversation to showcase the speed and performance. While it will be apparent in the app, I've pulled up the NPU on the performance monitor that's running on this device so that you can get an idea of what the efficiency looks like when you just send a regular prompt. So let's ask the model, "how are you doing?" And you can see that we get a response nearly instantly, and we also get a pretty generic LLM response, but that's expected because we are just chatting with a model. So now let's ask a question that again is just going to be using the model's general knowledge. And you can see that this model wasn't trained on the newest and latest information, so Snapdragon X isn't something it's intimately familiar with, but it does actually understand what an NPU is and that NPUs on devices can be very powerful for AI.
Demo: Chatting with a PDF Document
So here is a regular PDF that I've pulled off the internet that's about the Snapdragon X. And you can see this is a pretty complex PDF. We have multiple columns, embedded tables, even tiny text that is kind of all jumbled together, all of this giving us information about the Snapdragon X Elite. So let's upload this into AnythingLLM and then ask it about some questions or general information from this document.
To upload a document in AnythingLLM, it's as simple as just dragging it into the chat. This is going to embed the document privately on your device and will make it accessible to any other threads in this same workspace. You can reuse documents across workspaces. And while that was happening, you can see this tiny spike was actually the speed at which the whole document was embedded on the NPU. So now let's ask a question that's in this document that it would have more information on. So this is a pretty simple question of what is the name of the Snapdragon X Elite CPU. And we'll see that we have a spike in the usage of the NPU and then we start getting the result back, and it is correct. And it even gave us some additional facts about the CPU that were either related or adjacent to the question that we asked. Anytime an LLM is given context, we reference it as a citation that you can view here, and you can see the original context that was used to create that answer, just so that you always have a kind of truth or a basis for where LLM got an answer.
Data Connectors for Various Sources
But uploading documents via PDFs or files you have on your computer isn't the only way to get information into your workspace. AnythingLLM allows you to scrape specific websites, but also ships with many data connectors like connecting to GitHub, GitLab, pulling in a YouTube transcript, bulk scraping an entire website, and even connecting to Confluence. Data connectors are also customizable, so you can build the one that works best for you. But these are just some that we ship with out of the box to give you an idea of how much of anything AnythingLLM can use.
Unlocking New Capabilities with AI Agents
However, chatting with documents is a very, very simple use case of AnythingLLM, and AI agents actually unlock the ability for you to achieve much more with AnythingLLM on device, all powered on your NPU.
AnythingLLM ships with a bunch of agent skills that you can use right out of the box. However, building your own is very simple and you can bring it right into your desktop instance, or you can look on our community hub and download one that somebody may have already built that works just for you. Whatever you want to do, AnythingLLM can do it with AI agents. As you can see, our AI agent skills allow you to chat with your embedded documents, summarize documents, scrape websites in real time, generate files and charts, search the web, and even connect to an SQL database. Right now my web search provider is DuckDuckGo, but we support other search providers as well. I'm simply using DuckDuckGo because it's free to use.
Demo: Using an AI Agent for Real-Time Web Search
So now let's ask an agent-related question that would require real-time information to achieve. Agent chats work a little bit differently than chatting with the document, where they just begin with the text '@agent'. This kicks off an agent session where your LLM will have access to all those tools and any that you imported. And you'll see this little thought chain pop up as it starts to process our query and what tools it may be calling and the information associated. And you can see in pretty short order we were able to reach out to the internet, collect some kind of information, and use that in our LLM response and get an accurate response in real time. Now there's much more that you could do here. AI agents can chain many tools together and also with your custom tools be able to unlock a whole world of utility that otherwise has just been pretty difficult to get. And we've done all the hard work for you and it can run on your device privately today.
Conclusion
And so that's it for this short demo of AnythingLLM today. I hope you're as excited as we are about pioneering the way of running NPU-enabled models on device, fully privately, with all the tools that you need right at your fingertips. Thank you.