SQL Server Health Check

Find the SQL Server and Windows host issues behind slowdowns, failed jobs, and capacity pressure

A fixed-scope review for SQL Server estates running on Windows. DatAIbase reviews the technical evidence that affects performance, maintenance, backup reliability, storage, security, and day-to-day operation.

What the review covers

  • SQL Server instance, database, and workload evidence
  • Windows host capacity, drive space, services, and event-log indicators
  • Backup history, SQL Agent jobs, maintenance, indexes, and statistics
  • Severity-ranked findings with recommended next actions

Why teams book the review

Small technical issues can build into large production problems

SQL Server problems rarely arrive as one neat, isolated failure. A statistics problem can lead to poor estimates. Poor estimates can produce bad join choices, excessive reads, oversized memory grants, or spills to TempDB. TempDB pressure can turn into IO pressure. IO pressure can increase CPU time, extend query duration, and make blocking more visible. By the time users report slow screens, missed reporting windows, or timeouts, the system may be showing several issues at once.

The same pattern appears outside the query optimiser. A drive running short of free space can affect backups, data files, log growth, TempDB, and maintenance jobs. A disabled SQL Agent job can quietly remove an important safety net. A compatibility level left behind after an upgrade can leave performance features unused. A server that looked acceptable last year may not be ready for this year after another twelve months of deployments, data growth, and workload change.

The DatAIbase SQL Server Health Check is designed to give teams a practical technical baseline. It reviews SQL Server and Windows host evidence, separates high-priority issues from background noise, and produces a report that can be used for remediation planning, change discussions, and management decisions.

When to use it

Book the review when the system is still running, but confidence is starting to drop

The Health Check is not only for emergencies. It is useful when the team can see pressure building and needs an independent view of what the environment is showing.

Performance has drifted

Reports take longer, screens pause more often, batch jobs finish later, or the same workload now consumes more CPU, reads, memory, or elapsed time than it used to.

Storage is getting tight

Database, log, backup, or TempDB volumes are short of space, autogrowth events are becoming more frequent, or nobody is comfortable with how much headroom remains.

Maintenance is unreliable

SQL Agent jobs are failing, database maintenance is not completing reliably, or routine tasks such as CHECKDB and index maintenance are being missed.

Configuration drift has occurred over time

Instance and database settings were changed over several years and the team is no longer sure which settings are deliberate, inherited, temporary, or simply defaults.

A busy period or major change is coming

A seasonal peak, major deployment, migration, compatibility-level change, hardware refresh, or cloud move is approaching and the current baseline has not been reviewed recently.

DBA capacity is stretched

The internal team is busy with project work, BAU support, holidays, illness cover, or urgent incidents, and a structured external review would help prioritise the backlog.

Assessment scope

SQL Server and Windows host checks that produce useful technical evidence

The review focuses on evidence collected from SQL Server and the Windows host, so the report is based on the environment itself rather than guesswork.

1. Version, edition, and patch level

SQL Server version, edition, build number, database compatibility levels, support exposure, and upgrade indicators.

2. Instance configuration

Memory settings, MAXDOP, cost threshold for parallelism, backup compression default, optimise for ad hoc workloads, trace flags, and server-level configuration choices.

3. Windows host capacity

CPU, memory, operating system version, drive layout, drive free space, volume sizes, service state, and host-level indicators that can affect SQL Server operation.

4. Database settings and files

Recovery model, compatibility level, page verification, auto close, auto shrink, file sizes, autogrowth settings, log configuration, and database options.

5. TempDB configuration

TempDB file count, file sizes, growth settings, placement indicators, version-store pressure, spills where visible, and signs that TempDB is part of the workload problem.

6. Backup history

Full, differential, and log backup history from SQL Server, backup age, backup type coverage, backup size, compression, CHECKSUM usage where recorded, and obvious backup-chain issues.

7. SQL Agent and maintenance jobs

Job ownership, enabled and disabled jobs, recent failures, schedules, long-running jobs, maintenance routines, job notifications, and recurring SQL Agent error patterns.

8. Query Store and workload evidence

Query Store state, capture settings, high-resource queries, regressed queries where available, forced plans, failed forced plans, runtime history, and plan-cache indicators.

9. Waits, blocking, and resource pressure

Wait statistics, blocking evidence, deadlock information where available, memory grants, spills, IO indicators, CPU pressure, and workload hotspots.

10. Indexes and statistics

Missing, duplicate, unused, and overlapping indexes; stale or ineffective statistics; large tables with weak supporting indexes; and maintenance patterns that may create unnecessary data churn.

11. Security configuration

Sysadmin membership, SQL logins, orphaned users where detectable, database ownership, TRUSTWORTHY, xp_cmdshell, CLR, linked servers, and privileged access indicators.

12. High availability and SQL Server alerts

Availability Group, clustering, mirroring, and log shipping metadata where configured, plus SQL Agent Alerts, operators, Database Mail indicators, and job notifications.

Delivery process

A controlled review with a practical report at the end

The process is designed to be easy to approve and simple to run. The scope is agreed up front, evidence is collected from the SQL Server environment and Windows host, and DatAIbase produces a written report with a walkthrough of the findings.

How the review runs

  • Scope call: confirm the SQL Server instances, Windows hosts, timeframes, and access approach.
  • Evidence collection: gather SQL Server and Windows host data using an agreed, read-only collection approach.
  • Technical review: analyse configuration, capacity, maintenance, backup history, workload evidence, jobs, security settings, and host indicators.
  • Report production: group findings by severity, evidence, likely impact, and recommended action.
  • Walkthrough: review the findings, answer questions, and agree which actions should be considered first.

Report output

A clear technical report, prioritised for action

You receive a written assessment setting out each finding, the supporting evidence, the likely impact and the recommended next action. Findings are prioritised so your team can quickly see what needs attention first and where further investigation or change may be required.

Executive summary

Overall assessment of the environment
Highest-priority findings
Likely operational or business impact
Recommended immediate actions

Detailed findings

What was found and where
Evidence supporting the finding
Why it needs attention
Recommended action
Important dependencies or implementation considerations

Prioritised action plan

What should be addressed first
What can be scheduled through normal change control
What requires further investigation or application-owner input
What should be monitored rather than changed immediately

FAQ

SQL Server Health Check FAQ

What is a SQL Server Health Check?

A SQL Server Health Check is a structured assessment of the SQL Server environment and its Windows host. It looks for configuration, maintenance, capacity, recoverability, security and operational issues that may be affecting reliability or performance.

The output is a written assessment with prioritised findings, supporting evidence and recommended next actions.

What is the difference between a Health Check and a Performance Review?

A Health Check looks broadly across the SQL Server environment for issues that may need attention, even when there is no specific incident being investigated.
A SQL Server Performance Review starts with a known performance problem, such as slow queries, blocking, timeouts or workload instability, and investigates why it is happening.
If the main question is “What needs attention in this environment?”, choose the Health Check. If it is “Why is this workload slow?”, the Performance Review is usually the better fit.

Do you need direct access to our production SQL Server?

Not necessarily.

The collection method depends on the environment, your security requirements and the evidence required for the assessment. Access and data collection are agreed before the review begins.

Where possible, the aim is to collect the technical evidence required without introducing unnecessary access or disruption.

Will the Health Check change anything on our SQL Server?

No production changes are assumed as part of the Health Check.

The service is focused on assessment, evidence and recommendations. Any change resulting from the findings should be tested and introduced through your normal change-management process.

If DatAIbase assistance is required afterwards, that work can be scoped separately.

What does the Health Check review?

The exact evidence depends on the environment, but the assessment can include SQL Server instance and database configuration, Windows host resources, storage and drive capacity, backups, SQL Agent jobs, maintenance, indexes, statistics, Query Store and relevant security or operational indicators.

The purpose is not to run every possible check. It is to identify evidence that indicates risk, poor configuration or areas requiring further investigation.

Does the Health Check include drive space and Windows host checks?

Yes. SQL Server does not operate in isolation from the Windows host.

The review can include drive capacity, storage configuration, Windows services, host resource pressure and relevant operating-system evidence where those factors affect the SQL Server environment.

Does the Health Check include restore testing?

No. The Health Check can review backup history and relevant backup-chain indicators, but it does not claim that a database is recoverable simply because successful backups exist.

Restore testing or a more detailed recovery assessment can be scoped separately where required.

Does the Health Check review third-party monitoring tools?

Not as part of the standard assessment.

The Health Check focuses primarily on evidence from SQL Server and the Windows host. Existing monitoring data may be useful context where it is available, but reviewing the configuration, alerting strategy or effectiveness of a third-party monitoring platform would be a separate piece of work.

Is a Health Check only useful when something is already wrong?

No.

It can be used reactively when confidence in the environment has fallen, but it is also useful before an upgrade, migration, busy period, major release or other planned change.

The aim is to identify issues while there is still time to address them deliberately rather than during an incident.

What happens after the Health Check?

You receive a written assessment covering the important findings, the evidence behind them and the recommended next actions.

Findings are prioritised so your team can distinguish between issues that need prompt attention, changes that can be planned through normal change control, items requiring further investigation and observations that should simply be monitored.

If further assistance is required, implementation or deeper investigation can be scoped separately.