A2rchi does a quick level-up (v1.1.0)

Sweating in July and August in the hot Boston and Chicago summers the A2rchi team sets a new record fast release time. Less than one month has passed since the release of version 1.0.0 on July 16, that now version 1.1.0 is announced, which brings some noticeable improvements. They range from substantially sped-up inference calls for local GPU usage to allowing containers to use the faster direct host networking.

As usual there were various bug fixes and stability improvements, as well as a commensurate update of the users guide.


On a performance level: the vLLM package was interfaced as to offer another option than transformers to evaluate the LLM, when running inferences on a local server with GPUs. It turns out that for our test case it runs an order of magnitude faster while obtaining the same result. Configurations enable the user to control how many GPUs to run distributed on, or how much memory to allocate to vLLM, and more. A further speed improvement comes from a host mode option for the container networking that can be enabled via the configuration file. A2rchi now also “remembers” conversations more accurately. A2rchi’s internal summary of the conversation window is treated more carefully to best conserve the information provided.

As for integrations and interfaces, the Grafana monitoring has been upgraded, a vanilla Mattermost interface has been implemented, including now formatted history and context, amongst other changes. Additionally, there is now a scraper to snoop behind the CERN single-sign on (SSO) and a base SSO class to build other SSO scrapers, a Jira interface was added to store tickets from a specified URL and project into a vector database, and the Redmine service was integrated with PostgreSQL and Grafana monitoring.

Finally, from a developer standpoint, A2rchi has improved and uniformized logging in containers, templated prompts, and LLM outputs, which can now be found nicely organized in chain_input_output_log and stored in the container volume, ready for further debugging/studies.

Comments

Leave a Reply