Replacing matching entries in one column of a file by another column from a different fileMerge two files: two lines, partial line, two lines, partial line, etcFind common elements in a given column from two files and output the column values from each filecompare multiple files(more than two) with two different columnsReplace column in one file with column from another using awk?Joining columns from files if they contain a match in another columnMerging two files, one column at a timeColumn matching in separate filesExtract row if both column values appear in a single column from a separate fileJoining entries based off of column using awk/joinCompare two files by first column. Keep rows if matchingRecursively find and replace contents of one file using a key from another file
How does one intimidate enemies without having the capacity for violence?
Approximately how much travel time was saved by the opening of the Suez Canal in 1869?
How is it possible to have an ability score that is less than 3?
Accidentally leaked the solution to an assignment, what to do now? (I'm the prof)
Why doesn't H₄O²⁺ exist?
Convert two switches to a dual stack, and add outlet - possible here?
Java Casting: Java 11 throws LambdaConversionException while 1.8 does not
Are astronomers waiting to see something in an image from a gravitational lens that they've already seen in an adjacent image?
What would happen to a modern skyscraper if it rains micro blackholes?
Was any UN Security Council vote triple-vetoed?
What is the word for reserving something for yourself before others do?
Why can't I see bouncing of switch on oscilloscope screen?
What defenses are there against being summoned by the Gate spell?
I'm flying to France today and my passport expires in less than 2 months
Is it unprofessional to ask if a job posting on GlassDoor is real?
Codimension of non-flat locus
What's that red-plus icon near a text?
Watching something be written to a file live with tail
Important Resources for Dark Age Civilizations?
A newer friend of my brother's gave him a load of baseball cards that are supposedly extremely valuable. Is this a scam?
Do infinite dimensional systems make sense?
Unable to deploy metadata from Partner Developer scratch org because of extra fields
What are these boxed doors outside store fronts in New York?
Can I ask the recruiters in my resume to put the reason why I am rejected?
Replacing matching entries in one column of a file by another column from a different file
Merge two files: two lines, partial line, two lines, partial line, etcFind common elements in a given column from two files and output the column values from each filecompare multiple files(more than two) with two different columnsReplace column in one file with column from another using awk?Joining columns from files if they contain a match in another columnMerging two files, one column at a timeColumn matching in separate filesExtract row if both column values appear in a single column from a separate fileJoining entries based off of column using awk/joinCompare two files by first column. Keep rows if matchingRecursively find and replace contents of one file using a key from another file
.everyoneloves__top-leaderboard:empty,.everyoneloves__mid-leaderboard:empty,.everyoneloves__bot-mid-leaderboard:empty margin-bottom:0;
I have two tab-separated files which look as follows:
file1:
NC_008146.1 WP_011558474.1 1155234 1156286 44173
NC_008146.1 WP_011558475.1 1156298 1156807 12
NC_008146.1 WP_011558476.1 1156804 1157820 -3
NC_008705.1 WP_011558474.1 1159543 1160595 42748
NC_008705.1 WP_011558475.1 1160607 1161116 12
NC_008705.1 WP_011558476.1 1161113 1162129 -3
NC_009077.1 WP_011559727.1 2481079 2481633 8
NC_009077.1 WP_011854835.1 1163068 1164120 42559
NC_009077.1 WP_011854836.1 1164127 1164636 7
file2:
NC_008146.1 GCF_000014165.1_ASM1416v1_protein.faa
NC_008705.1 GCF_000015405.1_ASM1540v1_protein.faa
NC_009077.1 GCF_000016005.1_ASM1600v1_protein.faa
I want to match column 1 of file1 to file2 and replace itself with the respective column 2 entry of file 2.
The output would look like this:
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
awk
New contributor
add a comment |
I have two tab-separated files which look as follows:
file1:
NC_008146.1 WP_011558474.1 1155234 1156286 44173
NC_008146.1 WP_011558475.1 1156298 1156807 12
NC_008146.1 WP_011558476.1 1156804 1157820 -3
NC_008705.1 WP_011558474.1 1159543 1160595 42748
NC_008705.1 WP_011558475.1 1160607 1161116 12
NC_008705.1 WP_011558476.1 1161113 1162129 -3
NC_009077.1 WP_011559727.1 2481079 2481633 8
NC_009077.1 WP_011854835.1 1163068 1164120 42559
NC_009077.1 WP_011854836.1 1164127 1164636 7
file2:
NC_008146.1 GCF_000014165.1_ASM1416v1_protein.faa
NC_008705.1 GCF_000015405.1_ASM1540v1_protein.faa
NC_009077.1 GCF_000016005.1_ASM1600v1_protein.faa
I want to match column 1 of file1 to file2 and replace itself with the respective column 2 entry of file 2.
The output would look like this:
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
awk
New contributor
It looks like you might also be interested in our sister site: Bioinformatics.
– terdon♦
12 hours ago
Thank you for the link @terdon!
– BhushanDhamale
12 hours ago
add a comment |
I have two tab-separated files which look as follows:
file1:
NC_008146.1 WP_011558474.1 1155234 1156286 44173
NC_008146.1 WP_011558475.1 1156298 1156807 12
NC_008146.1 WP_011558476.1 1156804 1157820 -3
NC_008705.1 WP_011558474.1 1159543 1160595 42748
NC_008705.1 WP_011558475.1 1160607 1161116 12
NC_008705.1 WP_011558476.1 1161113 1162129 -3
NC_009077.1 WP_011559727.1 2481079 2481633 8
NC_009077.1 WP_011854835.1 1163068 1164120 42559
NC_009077.1 WP_011854836.1 1164127 1164636 7
file2:
NC_008146.1 GCF_000014165.1_ASM1416v1_protein.faa
NC_008705.1 GCF_000015405.1_ASM1540v1_protein.faa
NC_009077.1 GCF_000016005.1_ASM1600v1_protein.faa
I want to match column 1 of file1 to file2 and replace itself with the respective column 2 entry of file 2.
The output would look like this:
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
awk
New contributor
I have two tab-separated files which look as follows:
file1:
NC_008146.1 WP_011558474.1 1155234 1156286 44173
NC_008146.1 WP_011558475.1 1156298 1156807 12
NC_008146.1 WP_011558476.1 1156804 1157820 -3
NC_008705.1 WP_011558474.1 1159543 1160595 42748
NC_008705.1 WP_011558475.1 1160607 1161116 12
NC_008705.1 WP_011558476.1 1161113 1162129 -3
NC_009077.1 WP_011559727.1 2481079 2481633 8
NC_009077.1 WP_011854835.1 1163068 1164120 42559
NC_009077.1 WP_011854836.1 1164127 1164636 7
file2:
NC_008146.1 GCF_000014165.1_ASM1416v1_protein.faa
NC_008705.1 GCF_000015405.1_ASM1540v1_protein.faa
NC_009077.1 GCF_000016005.1_ASM1600v1_protein.faa
I want to match column 1 of file1 to file2 and replace itself with the respective column 2 entry of file 2.
The output would look like this:
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
awk
awk
New contributor
New contributor
edited 12 hours ago
Rui F Ribeiro
41.9k1483142
41.9k1483142
New contributor
asked 12 hours ago
BhushanDhamaleBhushanDhamale
1453
1453
New contributor
New contributor
It looks like you might also be interested in our sister site: Bioinformatics.
– terdon♦
12 hours ago
Thank you for the link @terdon!
– BhushanDhamale
12 hours ago
add a comment |
It looks like you might also be interested in our sister site: Bioinformatics.
– terdon♦
12 hours ago
Thank you for the link @terdon!
– BhushanDhamale
12 hours ago
It looks like you might also be interested in our sister site: Bioinformatics.
– terdon♦
12 hours ago
It looks like you might also be interested in our sister site: Bioinformatics.
– terdon♦
12 hours ago
Thank you for the link @terdon!
– BhushanDhamale
12 hours ago
Thank you for the link @terdon!
– BhushanDhamale
12 hours ago
add a comment |
2 Answers
2
active
oldest
votes
You can do this very easily with awk
:
$ awk 'NR==FNRa[$1]=$2; next$1=a[$1]; print' file2 file1
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
Or, since that looks like a tab-separated file:
$ awk -vOFS="t" 'NR==FNRa[$1]=$2; next$1=a[$1]; print' file2 file1
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
This assumes that every RefSeq (NC_*
) id in file1
has a corresponding entry in file2
.
Explanation
NR==FNR
: NR is the current line number, FNR is the line number of the current file. The two will be identical only while the 1st file (here,file2
) is being read.a[$1]=$2; next
: if this is the first file (see above), save the 2nd field in an array whose key is the 1st field. Then, move on to thenext
line. This ensures the next block isn't executed for the 1st file.$1=a[$1]; print
: now, in the second file, set the 1st field to whatever value was saved in the arraya
for the 1st field (so, the associated value fromfile2
) and print the resulting line.
1
NR == FNR
doesn't work correctly when the first file is empty. See this and the associated answer for a workaround
– iruvar
12 hours ago
1
@iruvar nothing will work well if the first file is empty, so I don't really see why that's relevant. The entire point here is to combine the data from the two files. If either file is empty, the whole exercise is pointless.
– terdon♦
12 hours ago
sorry I should have said in this particular casefile2
and notfile1
is empty. Sane behaviour whenfile2
is empty is to report the contents offile1
. The problem withNR == FNR
is that code associated with it executes on the contents offile1
whenfile2
is empty
– iruvar
12 hours ago
1
@iruvar there is no sane behavior here if either file is empty. That's what I'm saying :) So trying to make it deal with that case gracefully is pointless. And, in any case, when either file is empty here, nothing is printed. Which actually seems like the sanest approach, I'd rather get no data than wrong data.
– terdon♦
12 hours ago
add a comment |
No need for awk, assuming the files are sorted, you can use coreutils join:
join -o '2.2 1.2 1.3 1.4 1.5' file1 file2
Output:
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
If your files aren't, sorted, you can either sort them first (sort file1 > file1.sorted; sort file2 > file2.sorted
) and then use the command above, or, if your shell supports the <()
construct (bash does), you can do:
join -o '2.2 1.2 1.3 1.4 1.5' <(sort file1) <(sort file2)
add a comment |
Your Answer
StackExchange.ready(function()
var channelOptions =
tags: "".split(" "),
id: "106"
;
initTagRenderer("".split(" "), "".split(" "), channelOptions);
StackExchange.using("externalEditor", function()
// Have to fire editor after snippets, if snippets enabled
if (StackExchange.settings.snippets.snippetsEnabled)
StackExchange.using("snippets", function()
createEditor();
);
else
createEditor();
);
function createEditor()
StackExchange.prepareEditor(
heartbeatType: 'answer',
autoActivateHeartbeat: false,
convertImagesToLinks: false,
noModals: true,
showLowRepImageUploadWarning: true,
reputationToPostImages: null,
bindNavPrevention: true,
postfix: "",
imageUploader:
brandingHtml: "Powered by u003ca class="icon-imgur-white" href="https://imgur.com/"u003eu003c/au003e",
contentPolicyHtml: "User contributions licensed under u003ca href="https://creativecommons.org/licenses/by-sa/3.0/"u003ecc by-sa 3.0 with attribution requiredu003c/au003e u003ca href="https://stackoverflow.com/legal/content-policy"u003e(content policy)u003c/au003e",
allowUrls: true
,
onDemand: true,
discardSelector: ".discard-answer"
,immediatelyShowMarkdownHelp:true
);
);
BhushanDhamale is a new contributor. Be nice, and check out our Code of Conduct.
Sign up or log in
StackExchange.ready(function ()
StackExchange.helpers.onClickDraftSave('#login-link');
);
Sign up using Google
Sign up using Facebook
Sign up using Email and Password
Post as a guest
Required, but never shown
StackExchange.ready(
function ()
StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2funix.stackexchange.com%2fquestions%2f510709%2freplacing-matching-entries-in-one-column-of-a-file-by-another-column-from-a-diff%23new-answer', 'question_page');
);
Post as a guest
Required, but never shown
2 Answers
2
active
oldest
votes
2 Answers
2
active
oldest
votes
active
oldest
votes
active
oldest
votes
You can do this very easily with awk
:
$ awk 'NR==FNRa[$1]=$2; next$1=a[$1]; print' file2 file1
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
Or, since that looks like a tab-separated file:
$ awk -vOFS="t" 'NR==FNRa[$1]=$2; next$1=a[$1]; print' file2 file1
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
This assumes that every RefSeq (NC_*
) id in file1
has a corresponding entry in file2
.
Explanation
NR==FNR
: NR is the current line number, FNR is the line number of the current file. The two will be identical only while the 1st file (here,file2
) is being read.a[$1]=$2; next
: if this is the first file (see above), save the 2nd field in an array whose key is the 1st field. Then, move on to thenext
line. This ensures the next block isn't executed for the 1st file.$1=a[$1]; print
: now, in the second file, set the 1st field to whatever value was saved in the arraya
for the 1st field (so, the associated value fromfile2
) and print the resulting line.
1
NR == FNR
doesn't work correctly when the first file is empty. See this and the associated answer for a workaround
– iruvar
12 hours ago
1
@iruvar nothing will work well if the first file is empty, so I don't really see why that's relevant. The entire point here is to combine the data from the two files. If either file is empty, the whole exercise is pointless.
– terdon♦
12 hours ago
sorry I should have said in this particular casefile2
and notfile1
is empty. Sane behaviour whenfile2
is empty is to report the contents offile1
. The problem withNR == FNR
is that code associated with it executes on the contents offile1
whenfile2
is empty
– iruvar
12 hours ago
1
@iruvar there is no sane behavior here if either file is empty. That's what I'm saying :) So trying to make it deal with that case gracefully is pointless. And, in any case, when either file is empty here, nothing is printed. Which actually seems like the sanest approach, I'd rather get no data than wrong data.
– terdon♦
12 hours ago
add a comment |
You can do this very easily with awk
:
$ awk 'NR==FNRa[$1]=$2; next$1=a[$1]; print' file2 file1
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
Or, since that looks like a tab-separated file:
$ awk -vOFS="t" 'NR==FNRa[$1]=$2; next$1=a[$1]; print' file2 file1
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
This assumes that every RefSeq (NC_*
) id in file1
has a corresponding entry in file2
.
Explanation
NR==FNR
: NR is the current line number, FNR is the line number of the current file. The two will be identical only while the 1st file (here,file2
) is being read.a[$1]=$2; next
: if this is the first file (see above), save the 2nd field in an array whose key is the 1st field. Then, move on to thenext
line. This ensures the next block isn't executed for the 1st file.$1=a[$1]; print
: now, in the second file, set the 1st field to whatever value was saved in the arraya
for the 1st field (so, the associated value fromfile2
) and print the resulting line.
1
NR == FNR
doesn't work correctly when the first file is empty. See this and the associated answer for a workaround
– iruvar
12 hours ago
1
@iruvar nothing will work well if the first file is empty, so I don't really see why that's relevant. The entire point here is to combine the data from the two files. If either file is empty, the whole exercise is pointless.
– terdon♦
12 hours ago
sorry I should have said in this particular casefile2
and notfile1
is empty. Sane behaviour whenfile2
is empty is to report the contents offile1
. The problem withNR == FNR
is that code associated with it executes on the contents offile1
whenfile2
is empty
– iruvar
12 hours ago
1
@iruvar there is no sane behavior here if either file is empty. That's what I'm saying :) So trying to make it deal with that case gracefully is pointless. And, in any case, when either file is empty here, nothing is printed. Which actually seems like the sanest approach, I'd rather get no data than wrong data.
– terdon♦
12 hours ago
add a comment |
You can do this very easily with awk
:
$ awk 'NR==FNRa[$1]=$2; next$1=a[$1]; print' file2 file1
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
Or, since that looks like a tab-separated file:
$ awk -vOFS="t" 'NR==FNRa[$1]=$2; next$1=a[$1]; print' file2 file1
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
This assumes that every RefSeq (NC_*
) id in file1
has a corresponding entry in file2
.
Explanation
NR==FNR
: NR is the current line number, FNR is the line number of the current file. The two will be identical only while the 1st file (here,file2
) is being read.a[$1]=$2; next
: if this is the first file (see above), save the 2nd field in an array whose key is the 1st field. Then, move on to thenext
line. This ensures the next block isn't executed for the 1st file.$1=a[$1]; print
: now, in the second file, set the 1st field to whatever value was saved in the arraya
for the 1st field (so, the associated value fromfile2
) and print the resulting line.
You can do this very easily with awk
:
$ awk 'NR==FNRa[$1]=$2; next$1=a[$1]; print' file2 file1
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
Or, since that looks like a tab-separated file:
$ awk -vOFS="t" 'NR==FNRa[$1]=$2; next$1=a[$1]; print' file2 file1
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
This assumes that every RefSeq (NC_*
) id in file1
has a corresponding entry in file2
.
Explanation
NR==FNR
: NR is the current line number, FNR is the line number of the current file. The two will be identical only while the 1st file (here,file2
) is being read.a[$1]=$2; next
: if this is the first file (see above), save the 2nd field in an array whose key is the 1st field. Then, move on to thenext
line. This ensures the next block isn't executed for the 1st file.$1=a[$1]; print
: now, in the second file, set the 1st field to whatever value was saved in the arraya
for the 1st field (so, the associated value fromfile2
) and print the resulting line.
edited 12 hours ago
answered 12 hours ago
terdon♦terdon
133k33268448
133k33268448
1
NR == FNR
doesn't work correctly when the first file is empty. See this and the associated answer for a workaround
– iruvar
12 hours ago
1
@iruvar nothing will work well if the first file is empty, so I don't really see why that's relevant. The entire point here is to combine the data from the two files. If either file is empty, the whole exercise is pointless.
– terdon♦
12 hours ago
sorry I should have said in this particular casefile2
and notfile1
is empty. Sane behaviour whenfile2
is empty is to report the contents offile1
. The problem withNR == FNR
is that code associated with it executes on the contents offile1
whenfile2
is empty
– iruvar
12 hours ago
1
@iruvar there is no sane behavior here if either file is empty. That's what I'm saying :) So trying to make it deal with that case gracefully is pointless. And, in any case, when either file is empty here, nothing is printed. Which actually seems like the sanest approach, I'd rather get no data than wrong data.
– terdon♦
12 hours ago
add a comment |
1
NR == FNR
doesn't work correctly when the first file is empty. See this and the associated answer for a workaround
– iruvar
12 hours ago
1
@iruvar nothing will work well if the first file is empty, so I don't really see why that's relevant. The entire point here is to combine the data from the two files. If either file is empty, the whole exercise is pointless.
– terdon♦
12 hours ago
sorry I should have said in this particular casefile2
and notfile1
is empty. Sane behaviour whenfile2
is empty is to report the contents offile1
. The problem withNR == FNR
is that code associated with it executes on the contents offile1
whenfile2
is empty
– iruvar
12 hours ago
1
@iruvar there is no sane behavior here if either file is empty. That's what I'm saying :) So trying to make it deal with that case gracefully is pointless. And, in any case, when either file is empty here, nothing is printed. Which actually seems like the sanest approach, I'd rather get no data than wrong data.
– terdon♦
12 hours ago
1
1
NR == FNR
doesn't work correctly when the first file is empty. See this and the associated answer for a workaround– iruvar
12 hours ago
NR == FNR
doesn't work correctly when the first file is empty. See this and the associated answer for a workaround– iruvar
12 hours ago
1
1
@iruvar nothing will work well if the first file is empty, so I don't really see why that's relevant. The entire point here is to combine the data from the two files. If either file is empty, the whole exercise is pointless.
– terdon♦
12 hours ago
@iruvar nothing will work well if the first file is empty, so I don't really see why that's relevant. The entire point here is to combine the data from the two files. If either file is empty, the whole exercise is pointless.
– terdon♦
12 hours ago
sorry I should have said in this particular case
file2
and not file1
is empty. Sane behaviour when file2
is empty is to report the contents of file1
. The problem with NR == FNR
is that code associated with it executes on the contents of file1
when file2
is empty– iruvar
12 hours ago
sorry I should have said in this particular case
file2
and not file1
is empty. Sane behaviour when file2
is empty is to report the contents of file1
. The problem with NR == FNR
is that code associated with it executes on the contents of file1
when file2
is empty– iruvar
12 hours ago
1
1
@iruvar there is no sane behavior here if either file is empty. That's what I'm saying :) So trying to make it deal with that case gracefully is pointless. And, in any case, when either file is empty here, nothing is printed. Which actually seems like the sanest approach, I'd rather get no data than wrong data.
– terdon♦
12 hours ago
@iruvar there is no sane behavior here if either file is empty. That's what I'm saying :) So trying to make it deal with that case gracefully is pointless. And, in any case, when either file is empty here, nothing is printed. Which actually seems like the sanest approach, I'd rather get no data than wrong data.
– terdon♦
12 hours ago
add a comment |
No need for awk, assuming the files are sorted, you can use coreutils join:
join -o '2.2 1.2 1.3 1.4 1.5' file1 file2
Output:
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
If your files aren't, sorted, you can either sort them first (sort file1 > file1.sorted; sort file2 > file2.sorted
) and then use the command above, or, if your shell supports the <()
construct (bash does), you can do:
join -o '2.2 1.2 1.3 1.4 1.5' <(sort file1) <(sort file2)
add a comment |
No need for awk, assuming the files are sorted, you can use coreutils join:
join -o '2.2 1.2 1.3 1.4 1.5' file1 file2
Output:
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
If your files aren't, sorted, you can either sort them first (sort file1 > file1.sorted; sort file2 > file2.sorted
) and then use the command above, or, if your shell supports the <()
construct (bash does), you can do:
join -o '2.2 1.2 1.3 1.4 1.5' <(sort file1) <(sort file2)
add a comment |
No need for awk, assuming the files are sorted, you can use coreutils join:
join -o '2.2 1.2 1.3 1.4 1.5' file1 file2
Output:
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
If your files aren't, sorted, you can either sort them first (sort file1 > file1.sorted; sort file2 > file2.sorted
) and then use the command above, or, if your shell supports the <()
construct (bash does), you can do:
join -o '2.2 1.2 1.3 1.4 1.5' <(sort file1) <(sort file2)
No need for awk, assuming the files are sorted, you can use coreutils join:
join -o '2.2 1.2 1.3 1.4 1.5' file1 file2
Output:
GCF_000014165.1_ASM1416v1_protein.faa WP_011558474.1 1155234 1156286 44173
GCF_000014165.1_ASM1416v1_protein.faa WP_011558475.1 1156298 1156807 12
GCF_000014165.1_ASM1416v1_protein.faa WP_011558476.1 1156804 1157820 -3
GCF_000015405.1_ASM1540v1_protein.faa WP_011558474.1 1159543 1160595 42748
GCF_000015405.1_ASM1540v1_protein.faa WP_011558475.1 1160607 1161116 12
GCF_000015405.1_ASM1540v1_protein.faa WP_011558476.1 1161113 1162129 -3
GCF_000016005.1_ASM1600v1_protein.faa WP_011559727.1 2481079 2481633 8
GCF_000016005.1_ASM1600v1_protein.faa WP_011854835.1 1163068 1164120 42559
GCF_000016005.1_ASM1600v1_protein.faa WP_011854836.1 1164127 1164636 7
If your files aren't, sorted, you can either sort them first (sort file1 > file1.sorted; sort file2 > file2.sorted
) and then use the command above, or, if your shell supports the <()
construct (bash does), you can do:
join -o '2.2 1.2 1.3 1.4 1.5' <(sort file1) <(sort file2)
edited 12 hours ago
terdon♦
133k33268448
133k33268448
answered 12 hours ago
ThorThor
12.1k13762
12.1k13762
add a comment |
add a comment |
BhushanDhamale is a new contributor. Be nice, and check out our Code of Conduct.
BhushanDhamale is a new contributor. Be nice, and check out our Code of Conduct.
BhushanDhamale is a new contributor. Be nice, and check out our Code of Conduct.
BhushanDhamale is a new contributor. Be nice, and check out our Code of Conduct.
Thanks for contributing an answer to Unix & Linux Stack Exchange!
- Please be sure to answer the question. Provide details and share your research!
But avoid …
- Asking for help, clarification, or responding to other answers.
- Making statements based on opinion; back them up with references or personal experience.
To learn more, see our tips on writing great answers.
Sign up or log in
StackExchange.ready(function ()
StackExchange.helpers.onClickDraftSave('#login-link');
);
Sign up using Google
Sign up using Facebook
Sign up using Email and Password
Post as a guest
Required, but never shown
StackExchange.ready(
function ()
StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2funix.stackexchange.com%2fquestions%2f510709%2freplacing-matching-entries-in-one-column-of-a-file-by-another-column-from-a-diff%23new-answer', 'question_page');
);
Post as a guest
Required, but never shown
Sign up or log in
StackExchange.ready(function ()
StackExchange.helpers.onClickDraftSave('#login-link');
);
Sign up using Google
Sign up using Facebook
Sign up using Email and Password
Post as a guest
Required, but never shown
Sign up or log in
StackExchange.ready(function ()
StackExchange.helpers.onClickDraftSave('#login-link');
);
Sign up using Google
Sign up using Facebook
Sign up using Email and Password
Post as a guest
Required, but never shown
Sign up or log in
StackExchange.ready(function ()
StackExchange.helpers.onClickDraftSave('#login-link');
);
Sign up using Google
Sign up using Facebook
Sign up using Email and Password
Sign up using Google
Sign up using Facebook
Sign up using Email and Password
Post as a guest
Required, but never shown
Required, but never shown
Required, but never shown
Required, but never shown
Required, but never shown
Required, but never shown
Required, but never shown
Required, but never shown
Required, but never shown
It looks like you might also be interested in our sister site: Bioinformatics.
– terdon♦
12 hours ago
Thank you for the link @terdon!
– BhushanDhamale
12 hours ago