Personal data discovery that finds every place it sits, without copying any of it
Read-only scans of your databases, file shares and Microsoft 365 that keep counts and column names, never values. Indian identifiers checked by check digit. A record of processing that stays current, and the gaps on the record, not off it.
You cannot protect, retain or erase what you have not found
Data in places nobody listed
Exports on file shares, PAN in an invoice column, Aadhaar quoted in a notes field.
Tools built for another country
Scanners trained on foreign identifiers miss Indian ones, or flag every ten-digit number.
Discovery that copies the data
A scan that ships samples to a vendor's cloud creates the exposure it was meant to find.
A register that goes stale
A one-time data mapping exercise is out of date by the next quarter.
Data discovery, from every system to a record you can date
Every system, and how it is known
Scanned, declared by its owner in plain questions, cannot be scanned with a reason, or not yet covered. Before a scan, the credential's grants are read and any write privilege is reported.
Indian identifiers, checked before they are counted
Aadhaar by Verhoeff check, PAN, GSTIN, voter ID, passport, IFSC, UPI, ABHA, UAN and cards. Names by a model trained on Indian names that runs inside the platform. Found in free text too. Evidence is a sentence and a count, not a score.
A record of processing you can date
Whose data, purpose and basis, retention and sharing for every dataset. Versions carry a content hash and can be viewed as at any past date. Exports to Excel and Word.
Erasure verified at the source
Retention counted in rows past the period. Removal instructed to the owner, then checked by reading the source again; the owner's claim and the source's answer are recorded apart.
Vendors tiered by what they can reach
Every vendor, the personal data it can reach and its contract reviewed clause by clause. Assessments go as a link and are answered without an account.
Where one person's data is, in minutes
For a rights request or a breach, matching rows are counted in each connected source. The identifier is used in the query and stored only as a keyed hash.
The clocks start at awareness
CERT-In's six-hour report and the Board's clocks run from one timestamp. The scope comes from scanning the exposed copy, so the number of people affected is counted, not estimated.
What discovery reads
| Source | How it is read |
|---|---|
| PostgreSQL, MySQL, SQL Server, Oracle, SAP HANA | Read-only session; catalogue, samples and counts, paced to the source's latency |
| Snowflake, Databricks, BigQuery, MongoDB | Read-only service account |
| Windows and NFS file shares | CSV, TSV, Excel, PDF, Word and scanned images |
| Microsoft 365 and Google Workspace | SharePoint, OneDrive, Google Drive and mailboxes, read-only |
| Cloud storage and SaaS applications | S3, Azure Blob, Google Cloud Storage; HRMS, CRM and ticketing APIs |
| Anything else | Declared by its owner, or registered as cannot be scanned, with a reason |
Read-only
Discovery never writes to a source. Write privileges on the credential are flagged before a scan starts.
Metadata only
Samples are read in memory. Counts, column names and categories are kept. No value is stored.
Inside your infrastructure
No outbound calls except to endpoints you configure; keys from your key store.
Then act on what you found:Consent management software →Security assurance →
Data discovery, answered
Does the scanner copy our data?
No. Samples are read in memory inside your infrastructure. Only counts, column names and categories are kept.
Do we need discovery if we only have a few systems?
Often not. If personal data sits in a handful of systems you already know, a structured register may be enough. Discovery earns its place in larger and older estates.
How does it avoid false positives on Indian identifiers?
Identifiers are checked by format and, where the scheme has one, by check digit, such as Verhoeff for Aadhaar, so an invoice reference that looks like a PAN or an Aadhaar number is not counted as one.
What happens to systems it cannot scan?
They are registered with an owner, the reason, and the owner's answers to plain questions, so the record shows where its edges are.
Twenty questions. Ten minutes. Your exposure, by obligation.
No sign-up. Your answers stay in your browser. You see where liability sits and what would reduce it, including where software is not the answer.
See your own estate, read-only, in a day.
Wekalp is software for data fiduciaries and is not a Consent Manager registered with the Data Protection Board under section 6(9) of the Act. Nothing on this site is legal advice.