Delete duplicate records in SQL Server?

Question

Consider a column named EmployeeName table Employee. The goal is to delete repeated records, based on the EmployeeName field.

EmployeeName
------------
Anand
Anand
Anil
Dipak
Anil
Dipak
Dipak
Anil

Using one query, I want to delete the records which are repeated.

How can this be done with TSQL in SQL Server?

you could select the distinct values and their related IDs and delete those records whose IDs aren't in the already selected list? — DaeMoohn, Jul 23 '10 at 10:53
how did you accept the answer given by John Gibb, if table lacks of unique id? where is the `empId` column in your example used by John ? — armen, Oct 02 '13 at 11:01
If you don't have a unique ID column, or anything else meaningful to do an order by, you COULD also order by the employeename column... so your rn would be `row_number() over (partition by EmployeeName order by EmployeeName)`... this would pick an arbitrary single record for each name. — John Gibb, Nov 22 '13 at 04:20
Possible duplicate of [How can I remove duplicate rows?](https://stackoverflow.com/questions/18932/how-can-i-remove-duplicate-rows) — Erik, Dec 04 '17 at 22:17

score 243 · Accepted Answer · answered Jul 23 '10 at 15:22

243

You can do this with window functions. It will order the dupes by empId, and delete all but the first one.

delete x from (
  select *, rn=row_number() over (partition by EmployeeName order by empId)
  from Employee 
) x
where rn > 1;

Run it as a select to see what would be deleted:

select *
from (
  select *, rn=row_number() over (partition by EmployeeName order by empId)
  from Employee 
) x
where rn > 1;

answered Jul 23 '10 at 15:22

John Gibb

10,603
2
37
48

2

If you don't have a primary key, you can use `ORDER BY (SELECT NULL)` http://stackoverflow.com/a/4812038 – Arithmomaniac Jul 01 '16 at 16:57

StuartLC · Answer 2 · 2014-02-22T06:30:27.527

40

Assuming that your Employee table also has a unique column (ID in the example below), the following will work:

delete from Employee 
where ID not in
(
    select min(ID)
    from Employee 
    group by EmployeeName 
);

This will leave the version with the lowest ID in the table.

Edit
Re McGyver's comment - as of SQL 2012

MIN can be used with numeric, char, varchar, uniqueidentifier, or datetime columns, but not with bit columns

For 2008 R2 and earlier,

MIN can be used with numeric, char, varchar, or datetime columns, but not with bit columns (and it also doesn't work with GUID's)

For 2008R2 you'll need to cast the GUID to a type supported by MIN, e.g.

delete from GuidEmployees
where CAST(ID AS binary(16)) not in
(
    select min(CAST(ID AS binary(16)))
    from GuidEmployees
    group by EmployeeName 
);

SqlFiddle for various types in Sql 2008

SqlFiddle for various types in Sql 2012

edited Feb 22 '14 at 06:30

answered Jul 23 '10 at 11:07

StuartLC

104,537
17
209
285

Also, in Oracle, you could use "rowid" if there is no other unique id column. – Brandon Horsley Jul 23 '10 at 11:13
+1 Even if there were not an ID column, one could be added as an identity field. – Kyle B. Jul 23 '10 at 15:31
1

Excellent answer. Sharp and effective. Even if the table doesn't have an ID; it's better to include one to execute this method. – MiBol Nov 13 '18 at 17:07

Ben Cawley · Answer 3 · 2010-07-23T11:07:27.813

8

You could try something like the following:

delete T1
from MyTable T1, MyTable T2
where T1.dupField = T2.dupField
and T1.uniqueField > T2.uniqueField

(this assumes that you have an integer based unique field)

Personally though I'd say you were better off trying to correct the fact that duplicate entries are being added to the database before it occurs rather than as a post fix-it operation.

edited Jul 23 '10 at 11:07

answered Jul 23 '10 at 11:02

Ben Cawley

1,606
17
29

I donot have the unique field(ID) in my Table. How can i perform the operation then. – usr021986 Jul 24 '10 at 04:20

score 3 · Answer 4 · edited Aug 21 '13 at 07:24

DELETE
FROM MyTable
WHERE ID NOT IN (
     SELECT MAX(ID)
     FROM MyTable
     GROUP BY DuplicateColumn1, DuplicateColumn2, DuplicateColumn3)

WITH TempUsers (FirstName, LastName, duplicateRecordCount)
AS
(
    SELECT FirstName, LastName,
    ROW_NUMBER() OVER (PARTITIONBY FirstName, LastName ORDERBY FirstName) AS duplicateRecordCount
    FROM dbo.Users
)
DELETE
FROM TempUsers
WHERE duplicateRecordCount > 1

score 3 · Answer 5 · edited Aug 07 '14 at 21:10

3

WITH CTE AS
(
   SELECT EmployeeName, 
          ROW_NUMBER() OVER(PARTITION BY EmployeeName ORDER BY EmployeeName) AS R
   FROM employee_table
)
DELETE CTE WHERE R > 1;

The magic of common table expressions.

edited Aug 07 '14 at 21:10

dfrevert

366
5
17

answered Jul 25 '10 at 11:30

Mostafa Elmoghazi

2,124
1
21
27

SubPortal / a_horse_with_no_name - shouldn't this be selecting from an actual table? Also, ROW_NUMBER should be ROW_NUMBER() because it's a function, correct? – JustBeingHelpful Feb 17 '14 at 07:23

score 1 · Answer 6 · edited Oct 02 '13 at 10:49

1

Try

DELETE
FROM employee
WHERE rowid NOT IN (SELECT MAX(rowid) FROM employee
GROUP BY EmployeeName);

edited Oct 02 '13 at 10:49

Tamil Selvan C

19,913
12
49
70

answered Oct 02 '13 at 10:32

Anurag Garg

11
1

score 1 · Answer 7 · answered Sep 28 '16 at 06:57

If you're looking for a way to remove duplicates, yet you have a foreign key pointing to the table with duplicates, you could take the following approach using a slow yet effective cursor.

It will relocate the duplicate keys on the foreign key table.

create table #properOlvChangeCodes(
    id int not null,
    name nvarchar(max) not null
)

DECLARE @name VARCHAR(MAX);
DECLARE @id INT;
DECLARE @newid INT;
DECLARE @oldid INT;

DECLARE OLVTRCCursor CURSOR FOR SELECT id, name FROM Sales_OrderLineVersionChangeReasonCode; 
OPEN OLVTRCCursor;
FETCH NEXT FROM OLVTRCCursor INTO @id, @name;
WHILE @@FETCH_STATUS = 0  
BEGIN  
        -- determine if it should be replaced (is already in temptable with name)
        if(exists(select * from #properOlvChangeCodes where Name=@name)) begin
            -- if it is, finds its id
            Select  top 1 @newid = id
            from    Sales_OrderLineVersionChangeReasonCode
            where   Name = @name

            -- replace terminationreasoncodeid in olv for the new terminationreasoncodeid
            update Sales_OrderLineVersion set ChangeReasonCodeId = @newid where ChangeReasonCodeId = @id

            -- delete the record from the terminationreasoncode
            delete from Sales_OrderLineVersionChangeReasonCode where Id = @id
        end else begin
            -- insert into temp table if new
            insert into #properOlvChangeCodes(Id, name)
            values(@id, @name)
        end

        FETCH NEXT FROM OLVTRCCursor INTO @id, @name;
END;
CLOSE OLVTRCCursor;
DEALLOCATE OLVTRCCursor;

drop table #properOlvChangeCodes

score 0 · Answer 8 · answered Oct 21 '19 at 13:59

0

delete from person 
where ID not in
(
        select t.id from 
        (select min(ID) as id from person 
         group by email 
        ) as t
);

answered Oct 21 '19 at 13:59

ohsoifelse

681
7
6

Jithin Shaji · Answer 9 · 2016-09-28T10:44:43.070

Please see the below way of deletion too.

Declare @Employee table (EmployeeName varchar(10))

Insert into @Employee values 
('Anand'),('Anand'),('Anil'),('Dipak'),
('Anil'),('Dipak'),('Dipak'),('Anil')

Select * from @Employee

Created a sample table named @Employee and loaded it with given data.

Delete  aliasName from (
Select  *,
        ROW_NUMBER() over (Partition by EmployeeName order by EmployeeName) as rowNumber
From    @Employee) aliasName 
Where   rowNumber > 1

Select * from @Employee

Result:

I know, this is asked six years ago, posting just incase it is helpful for anyone.

score -1 · Answer 10 · answered Mar 13 '18 at 19:45

Here's a nice way of deduplicating records in a table that has an identity column based on a desired primary key that you can define at runtime. Before I start I'll populate a sample data set to work with using the following code:

if exists (select 1 from sys.all_objects where type='u' and name='_original')
drop table _original

declare @startyear int = 2017
declare @endyear int = 2018
declare @iterator int = 1
declare @income money = cast((SELECT round(RAND()*(5000-4990)+4990 , 2)) as money)
declare @salesrepid int = cast(floor(rand()*(9100-9000)+9000) as varchar(4))
create table #original (rowid int identity, monthyear varchar(max), salesrepid int, sale money)
while @iterator<=50000 begin
insert #original 
select (Select cast(floor(rand()*(@endyear-@startyear)+@startyear) as varchar(4))+'-'+ cast(floor(rand()*(13-1)+1) as varchar(2)) ),  @salesrepid , @income
set  @salesrepid  = cast(floor(rand()*(9100-9000)+9000) as varchar(4))
set @income = cast((SELECT round(RAND()*(5000-4990)+4990 , 2)) as money)
set @iterator=@iterator+1
end  
update #original
set monthyear=replace(monthyear, '-', '-0') where  len(monthyear)=6

select * into _original from #original

Next I'll create a Type called ColumnNames:

create type ColumnNames AS table   
(Columnnames varchar(max))

Finally I will create a stored proc with the following 3 caveats: 1. The proc will take a required parameter @tablename that defines the name of the table you are deleting from in your database. 2. The proc has an optional parameter @columns that you can use to define the fields that make up the desired primary key that you are deleting against. If this field is left blank, it is assumed that all the fields besides the identity column make up the desired primary key. 3. When duplicate records are deleted, the record with the lowest value in it's identity column will be maintained.

Here is my delete_dupes stored proc:

 create proc delete_dupes (@tablename varchar(max), @columns columnnames readonly) 
 as
 begin

declare @table table (iterator int, name varchar(max), is_identity int)
declare @tablepartition table (idx int identity, type varchar(max), value varchar(max))
declare @partitionby varchar(max)  
declare @iterator int= 1 


if exists (select 1 from @columns)  begin
declare @columns1 table (iterator int, columnnames varchar(max))
insert @columns1
select 1, columnnames from @columns
set @partitionby = (select distinct 
                substring((Select ', '+t1.columnnames 
                From @columns1 t1
                Where T1.iterator = T2.iterator
                ORDER BY T1.iterator
                For XML PATH ('')),2, 1000)  partition
From @columns1 T2 )

end

insert @table 
select 1, a.name, is_identity from sys.all_columns a join sys.all_objects b on a.object_id=b.object_id
where b.name = @tablename  

declare @identity varchar(max)= (select name from @table where is_identity=1)

while @iterator>=0 begin 
insert @tablepartition
Select          distinct case when @iterator=1 then 'order by' else 'over (partition by' end , 
                substring((Select ', '+t1.name 
                From @table t1
                Where T1.iterator = T2.iterator and is_identity=@iterator
                ORDER BY T1.iterator
                For XML PATH ('')),2, 5000)  partition
From @table T2
set @iterator=@iterator-1
end 

declare @originalpartition varchar(max)

if @partitionby is null begin
select @originalpartition  = replace(b.value+','+a.type+a.value ,'over (partition by','')  from @tablepartition a cross join @tablepartition b where a.idx=2 and b.idx=1
select @partitionby = a.type+a.value+' '+b.type+a.value+','+b.value+') rownum' from @tablepartition a cross join @tablepartition b where a.idx=2 and b.idx=1
 end
 else
 begin
 select @originalpartition=b.value +','+ @partitionby from @tablepartition a cross join @tablepartition b where a.idx=2 and b.idx=1
 set @partitionby = (select 'OVER (partition by'+ @partitionby  + ' ORDER BY'+ @partitionby + ','+b.value +') rownum'
 from @tablepartition a cross join @tablepartition b where a.idx=2 and b.idx=1)
 end


exec('select row_number() ' + @partitionby +', '+@originalpartition+' into ##temp from '+ @tablename+'')


exec(
'delete a from _original a 
left join ##temp b on a.'+@identity+'=b.'+@identity+' and rownum=1  
where b.rownum is null')

drop table ##temp

end

Once this is complied, you can delete all your duplicate records by running the proc. To delete dupes without defining a desired primary key use this call:

exec delete_dupes '_original'

To delete dupes based on a defined desired primary key use this call:

declare @table1 as columnnames
insert @table1
values ('salesrepid'),('sale')
exec delete_dupes '_original' , @table1

Delete duplicate records in SQL Server?

10 Answers10

Linked

Related