building an extensible genai copilot what we learned
Building an Extensible GenAI Copilot: What We Learned
September 30, 2024
Rajat Tiwari
Senior Software Engineer II
Working through the complexities of developing an internal copilot helped us push the boundaries of what we believed possible with GenAI.
Our generative AI (GenAI) journey began with a single use case: How could we make it easier for our customers to navigate the vast landscape of documentation and features within our platform?
It’s crucial that our users can quickly access and understand the product information they need to use our enterprise platform. So, our product team tasked our development team with developing a solution to simplify this process and enhance the overall user experience. A few months ago, we launched Rafay Copilot, a GenAI-driven bot designed to do just that.
However, these challenges provided us with invaluable insights. As we worked through the complexities of creating Rafay Copilot, we began to see the broader potential of GenAI. The problems we solved and our breakthroughs led us to realize that what we had developed could be expanded far beyond its original scope and into other areas of our platform.
Rafay’s Cloud Automation Platform provides a solution for platform teams that wish to build automated self-service cloud infrastructure workflows.
Defining the Copilot Architecture
GenAI has the potential to apply to many use cases in our product, so a major goal while defining the Rafay Copilot architecture was to ensure that it could be easily extended to support other use cases. Such an architecture should be flexible and scalable to support more advanced use cases such as agents or other types of copilots.
We used the LangChain framework to chain the requests before sending them to the large language models (LLMs). Qdrant serves as our vector database, efficiently storing and retrieving text embeddings. Retrieval-augmented generation (RAG) allows the chatbot to access and retrieve specific private data, such as company documentation or proprietary knowledge, to provide more accurate contextual responses.
API Gateway and AI Service Layer
All incoming requests to Rafay Copilot go through a centralized API gateway, which we call rafay-hub. This gateway’s primary responsibility is to handle authentication and standardize API requests for our upstream services. Once authenticated and standardized, the request is forwarded to the AI service, which serves as a proxy for all our agents.
The AI service’s primary role is to decide which agent service should handle the request. We’ve designed this so that it’s easy for us to extend the system in the future.
Agent Services and LLM Integration
Once the AI service determines the appropriate agent, the request is sent to the corresponding agent service. This agent interacts directly with the underlying LLMs. The prompts that guide these interactions are stored in a Kubernetes (K8s) config map, allowing us to adjust them quickly based on system performance or user feedback.
Vector Database and Observability
To provide accurate information from our documentation, we’ve implemented a Kubernetes cron job that regularly pulls data from our GitHub repository, where all the docs are stored in Markdown format. This job processes the data and stores it in Qdrant, our vector database, making it easy for the chatbot to retrieve relevant information quickly.
Overcoming Challenges
Even though the architecture diagram above may make it appear that the process was straightforward, we faced multiple hurdles while building Rafay Copilot. Development posed significant challenges:
- Steep learning curve: For any organization beginning its AI journey, understanding the AI landscape presents a steep learning curve.
- Evaluating LLM options: With the rise of Generative AI, an overwhelming number of LLMs are available, and deciding which LLM is best for your business is challenging.
- Governance:
- Cost management: Having governance in place to monitor and control costs is crucial.
- Guardrails: Security is critical to ensure sensitive data isn't leaked when interacting with LLMs.
- Secrets management: Managing API keys in an enterprise setting can be complex.
- Prompt evaluation: Even small changes in wording can lead to significantly different responses, making prompt evaluation crucial.
- Observability: Strong observability is required for AI apps, and integrating tools for ongoing maintenance takes time.
Conclusion
Building GenAI applications may sound easy. However, deploying enterprise-grade GenAI applications in production has many challenges, including finding the right LLM and maintaining observability. We used the insights gained from building Rafay Copilot to develop GenAI Playgrounds, a rapid prototyping IDE for developers to build GenAI applications.