Skip to content
Case Study · Enterprise Search

Enterprise Search Software for a Regulated Firm: Permission-Aware Answers Across SharePoint and File Shares

A regulated mid-market firm moved off Glean to an Onyx-based private AI stack running inside its own infrastructure, cutting search running cost by 85% while keeping every access rule enforced at query time.

Services
Glean Migration, Self-Hosted Enterprise Search, Onyx Deployment, Permission-Aware Retrieval
Client
Confidential, regulated mid-market firm
Date
2024–2025
Scale
3.5M documents · SharePoint · file shares · wiki · ticketing
Enterprise Search Software for a Regulated Firm: Permission-Aware Answers Across SharePoint and File Shares
85% lower
Search running cost after moving off Glean to a self-hosted stack
3.5M
Documents indexed across SharePoint, file shares, the wiki and ticketing
1,000/day
Searches run against the corpus, firm-wide
Enforced
Entitlements filtered at query time against the source systems
The Brief

Replacing Glean With a Private Stack the Firm Owns

The firm was already running Glean. It had solved the first problem, which was that knowledge sat scattered across a decade of SharePoint sites, departmental file shares, a wiki two teams maintained and a ticketing system. People could search across all of it, and they did.

What did not work was the shape of the arrangement. The corpus lived in a vendor's cloud, the permission model was a copy rather than the source, and the licence was priced per seat across an organisation where heavy users were always going to be a minority. At renewal those three things were harder to defend than they had been at signature, and the question became whether the same capability could run on infrastructure the firm already operated.

The migration, in one line: the same search behaviour, moved from a per-seat SaaS licence with the corpus in a vendor cloud, to an Onyx stack inside the firm's own boundary at 85% lower running cost, with entitlements enforced against the source systems rather than a mirrored copy.
On this page
Stack
Onyx Permission-aware RAG SharePoint connector SSO Self-hosted LLM On-prem
Challenges

Four Problems Standing in the Way

01

Per-Seat Licensing Across a Firm Where Use Was Uneven

Glean is licensed per seat across the whole organisation
Heavy users were always going to be a minority of staff
Cost scaled with headcount rather than with value delivered
02

The Corpus Sat in a Vendor's Cloud

Regulatory obligations cover where data is processed and stored
Every document indexed meant more of the firm outside its boundary
Defensible at signature, harder to defend as the corpus grew
03

Entitlements Were a Copy, Not the Source

Hosted search rebuilds its own version of your permission model
It stays roughly in step with the source systems, and roughly is the problem
One stale rule surfaces a document to the wrong person, which is reportable
04

Coverage Stopped Short of the Older Systems

Most institutional knowledge sat in ageing file shares and the wiki
Connector catalogues are built around modern SaaS tools first
Extending coverage meant waiting on someone else's roadmap
Production

The Solution

We deployed Onyx — the open-source enterprise search and AI assistant platform — inside the firm's own environment, connected to the systems the work actually lives in.

01

Permission-Aware Retrieval

Onyx syncs each source system's access control lists, so results filter against the individual's real entitlements at query time.

02

Connectors to Systems in Use

SharePoint, file shares, the wiki and ticketing indexed on a schedule, so answers reflect current documents.

03

Answers With Citations

RAG over the indexed corpus, every answer citing the sources it drew on — an uncited answer is not usable here.

04

Models Under Firm Control

Inference runs against models the firm chooses, hosted in its own environment, so no content is sent to an external API.

05

Rollout by Team

Deployed to one department first, its feedback shaping connector scope before the wider rollout.

Identity & access

SSO against the existing identity provider, group membership driving entitlements
Access control lists re-synced on each indexing run so permission changes propagate
Two people running an identical query correctly see different results

Indexing & audit

Connectors per source with scheduled re-indexing; retention scoped in or out deliberately
Every query and answer logged — a requirement in a regulated firm, not a feature
The log doubles as a map of what people ask and where the corpus has gaps
Impact

What It Moved for the Business

One Place to Ask

SharePoint, file shares, the wiki and ticketing are searchable together, so nobody has to guess which system holds the answer.

Entitlements Survived

Access rules are enforced against the source systems at query time rather than duplicated into a model that drifts.

Answers Are Checkable

Every response cites the documents behind it, so a user can open the source and verify before acting.

The Corpus Never Left

Documents, index and inference all remain inside the firm's network boundary, which is what allowed the project to proceed at all.

Adoption That Held

Around 1,000 searches a day across the firm, which is the bar an earlier internal search rollout never reached.

Conclusion

The firm kept its entitlement model intact, kept its corpus inside its own boundary, and gave staff a single place to ask a question in plain language.

Onyx's admin analytics record every query and answer, so the firm now has a running measure of adoption, of what people ask most, and of where the corpus has gaps — the questions that return poor answers are a direct map of the documentation not yet written. If your documents are spread across systems that each search only themselves, and access rules mean you cannot simply index everything into a hosted product, that is the position this build was designed for.

Documents Scattered Across Systems That Only Search Themselves?

Bring a list of the systems you would want indexed and the access rules that govern them. In 30 minutes we'll tell you whether a self-hosted, permission-aware search fits — and what it takes to stand up.

Discuss your AI project Discuss your project