Who it’s for / IT & Infrastructure
Grotabyte for IT and Infrastructure
Compliance signs the contract; you are the one who gets paged. The question that matters to you is not what the archive can do in a demo, but what it costs to keep alive at three in the morning, and how you find out when capture has quietly stopped.
Rated 5 out of 5. “Ten-year holds, answered by self-service.” — College of DuPage
Also for: CIO · CTO · VP IT · IT Director · Systems Administrator
What you are accountable for
You are accountable for the archive actually running — capture that does not fail silently, storage that does not surprise you, and a platform that does not turn into another cluster with your name on it.
What makes that hard
Every new platform is a new on-call rotation
Search clusters need nodes, shard rebalances, version upgrades and someone who understands them at two in the morning. The archive is bought for a compliance reason and operated on your headcount.
Capture fails quietly
A journal mailbox goes silent, a connector's watermark stops advancing, an OAuth grant expires. Nothing breaks visibly. You learn about the gap months later, in a matter, when someone asks where a custodian's mail went.
The decision you will be asked to reverse
You choose cloud, and eighteen months later a regulator, a board or an acquisition means it has to run on your own hardware. If that is a different product, a migration and a re-index, the first decision was a trap.
Retrieval tickets land on your queue
Every 'can you find an email from 2019' request routes to IT, because the archive was built for legal and nobody else can use it.
What Grotabyte gives you
No search cluster to operate
Search runs on embedded engines over object storage — there is no separate search tier to size, shard, rebalance or upgrade. The thing you are being asked to adopt is not a distributed database with a UI on top of it.
Failures that announce themselves
Stalled collection watermarks and quiet journal mailboxes are graded as alerts rather than left to be discovered in a matter. A collection schedule that cannot survive the source vendor's own retention window is refused by name at configuration time, before the gap opens.
One build, three deployments
On-premise, in the cloud, or run by us — the same build, selected by one environment variable. On-premise runs against your own filesystem, NAS or SAN, with record keys that never leave the machine and no outbound dependency required to search or produce.
Every mailbox you have, cloud, on-prem or long-dead
Microsoft 365, Google Workspace, on-premise Exchange over EWS including journal mailboxes, any IMAP or POP3 mailbox, S3, R2, SFTP, NAS shares, and browser upload. Legacy material comes in as PST, OST, EML, MSG, MBOX, EMLX and DXL — and PST is written natively as well as read.
Access you configure once
Passwordless sign-in and per-tenant OIDC routed by DNS-verified domain, with 25 permissions across 10 roles. Give staff self-service search of their own history and the retrieval tickets stop arriving at your desk.
One checkable thing
College of DuPage runs ten-year litigation holds, answers public-records requests faster, and lets staff search their own archived history rather than raising a ticket with IT.
Questions this role asks
What do we actually have to operate?
The application and object storage. There is no search cluster: the engines are embedded and run over object storage, so there are no nodes to size, no shards to rebalance and no separate search tier to keep on a version. In an on-premise deployment the object storage can be your own filesystem, NAS or SAN.
Can we start in the cloud and move on-premise later, or the other way round?
Yes — it is the same build in both cases, selected by a single environment variable, not a different product with a shared brand. On-premise deployments run against your own storage with record keys that never leave the machine and no outbound dependency needed to search or produce, which is usually the reason the question comes up in the first place.
How do we know capture hasn't quietly stopped?
Collection health is graded, not merely logged. A watermark that stops advancing and a journal mailbox that has gone quiet both raise alerts, and a schedule that could not survive the source vendor's retention window is rejected when you configure it, with the reason named, rather than accepted and left to leave a hole.