Text Extraction and retrieval

Question

John il 24 Ott 2017

0
Link

Link diretto a questa domanda

https://it.mathworks.com/matlabcentral/answers/363069-text-extraction-and-retrieval

Modificato: shilpa patil il 23 Set 2019

 <P ID=1>
A LITTLE BLACK BIRD.
</P>
 <P ID=2>
Story about a bird, 
(1811)
</P>
 <P ID=3>
Part 1.
</P>

As I am new to text extraction, I need help in;

Writing a code to count the delimiters (</P>)
Remove all punctuation
Break the text into individual documents at each delimiter, knowing that ID=1 refers to document 1, ID=2 refers to document 2. etc

0 Commenti
Mostra -2 commenti meno recentiNascondi -2 commenti meno recenti

Accedi per commentare.

Accedi per rispondere a questa domanda.

Answer 1

Akira Agata il 25 Ott 2017

1
Link

Link diretto a questa risposta

https://it.mathworks.com/matlabcentral/answers/363069-text-extraction-and-retrieval#answer_287567

Apri in MATLAB Online

Just tried to make a script to do that. Here is the result (assuming the maximum ID = 10).

% Read your text file
fid = fopen('yourText.txt');
C = textscan(fid,'%s','TextType','string','Delimiter','\n','EndOfLine','\r\n');
C = C{1};
fclose(fid);
% 1. Count the delimiters '</P>'
idx = strfind(C,'</P>');
n = nnz(cellfun(@(x) ~isempty(x), idx));
% 2. Remove all punctuation
C2 = regexprep(C,'[.,!?:;]','');
% 3. Break the text into individual documents at each delimiter
idx2 = find(strcmp(C,'</P>'));
for kk = 1:10
  str = ['<P ID=',num2str(kk),'>'];
  idx_s = find(strcmp(C,str));
  if ~isempty(idx_s)
      idx_e = idx2(find(idx2>idx_s,1));
      fileName = ['document',num2str(kk),'.txt'];
      fid = fopen(fileName,'w');
      fprintf(fid,'%s\r\n',C(idx_s:idx_e));
      fclose(fid);
  end    
end

6 Commenti
Mostra 4 commenti meno recentiNascondi 4 commenti meno recenti

Akira Agata il 30 Ott 2017

Modificato: Akira Agata il 30 Ott 2017

Apri in MATLAB Online

Thanks for your reply. I've just made a script to do the items 1~3, as follows. I hope this will help you somehow.

Regarding your last question ("count the number of documents each word appear in"), I think you can do that by combining the following script with my previous one.

% Read your text file
fid = fopen('yourText.txt');
C = textscan(fid,'%s','TextType','string','Delimiter','\n','EndOfLine','\r\n');
C = C{1};
fclose(fid);
C = regexprep(C,'<[\w \=\/]+>',''); % Remove tags
C = regexprep(C,'[.,!?:;()]','');   % Remove punctuation and brackets
C = regexprep(C,'[0-9]+','');       % Remove numbers
C = lower(C);                       % Convert to lower case
% Extract every words
words = regexp(C,'[a-z\-]+','match');
words = [words{:}];
% (1) Count total number of words
numOfWords = numel(words); % --> 9
% (2) Count the total number of distinct words
numOfDistWords = numel(unique(words)); % --> 7
% (3) Find the number of times each word is used in the original text
wordList = unique(words);
wordCount = arrayfun(@(x) nnz(strcmp(x,words)), wordList);
% Show the result
figure
bar(wordCount)
xticklabels(wordList)

John il 7 Nov 2017

Apri in MATLAB Online

Thanks. I am stuck running the counter.

for kk = 1:n
  str = ['<p id=',num2str(kk),'>'];
  idx_s = find(strcmp(C,str));
  if ~isempty(idx_s)
      idx_e = idx2(find(idx2>idx_s,1));
      Doc=C(idx_s:idx_e); %May need to remove tags later
      Doc = regexp(Doc,'[a-z0-9\-]+','match');
      Doc = [Doc{:}];
      Unique_Doc_count = arrayfun(@(x) nnz(strcmp(x,Doc)), Unique);
      Unique_Doc_freq=[Unique;Unique_Doc_count];
  end    
end

I want to search if the elements in string array 'Unique' exist in 'Doc'. I got results in 'Unique_Doc_count' as the number of their occurrences but I need just 1 or 0 values (exist) or (not exist). The aim is to loop 'kk' over multiple documents and find the number of documents that contain each word in 'Unique'. Not even number of times the word occurs, but number of documents it appears in.

Accedi per commentare.

Answer 2

Cedric il 26 Ott 2017

2
Link

Link diretto a questa risposta

https://it.mathworks.com/matlabcentral/answers/363069-text-extraction-and-retrieval#answer_287987

Apri in MATLAB Online

Here is another approach based on pattern matching:

 >> data = regexp(fileread('data.txt'), '(?<=<P[^>]+>\s*)[\w ]+', 'match' )
 data =
  1×3 cell array
    {'A LITTLE BLACK BIRD'}    {'Story about a bird'}    {'Part 1'}

if you don't need the IDs (e.g. if in any case they will go from 1 to the number of P tags), you are done.

If you needed the IDs, you could get both IDs and content as follows:

 >> data = regexp(fileread('data.txt'), '<P ID=(\d+)>\s*([\w ]+)', 'tokens' ) ;
    data = vertcat( data{:} ) ;
    ids  = str2double( data(:,1) )
    data = data(:,2)
 ids =
     1
     2
     3
 data =
  3×1 cell array
    {'A LITTLE BLACK BIRD'}
    {'Story about a bird' }
    {'Part 1'             }

6 Commenti
Mostra 4 commenti meno recentiNascondi 4 commenti meno recenti

John il 7 Nov 2017

Apri in MATLAB Online

Thanks. I am stuck running a counter.

for kk = 1:n
  str = ['<p id=',num2str(kk),'>'];
  idx_s = find(strcmp(C,str));
  if ~isempty(idx_s)
      idx_e = idx2(find(idx2>idx_s,1));
      Doc=C(idx_s:idx_e); %May need to remove tags later
      Doc = regexp(Doc,'[a-z0-9\-]+','match');
      Doc = [Doc{:}];
      Unique_Doc_count = arrayfun(@(x) nnz(strcmp(x,Doc)), Unique);
      Unique_Doc_freq=[Unique;Unique_Doc_count];
  end    
end

I want to search if the elements in string array 'Unique' exist in 'Doc'. I got results in 'Unique_Doc_count' as the number of their occurrences but I need just 1 or 0 values (exist) or (not exist). The aim is to loop 'kk' over multiple documents and find the number of documents that contain each word in 'Unique'. Not even number of times the word occurs, but number of documents it appears in.

Cedric il 9 Nov 2017

Modificato: Cedric il 9 Nov 2017

Apri in MATLAB Online

If you have a count per document, finding the number of documents a keyword is in is easy:

 counts = [7, 0 ,3] ;
 hasKey = counts > 0 ;        % [1,0,1]
 nDocs  = sum( hasKey ) ;     % 2

Accedi per commentare.

Answer 3

Christopher Creutzig il 2 Nov 2017

0
Link

Link diretto a questa risposta

https://it.mathworks.com/matlabcentral/answers/363069-text-extraction-and-retrieval#answer_289028

Modificato: Christopher Creutzig il 2 Nov 2017

Apri in MATLAB Online

It's probably easiest to split the text and then check the number of splits created to count, using string functions:

str = extractFileText('file.txt');
paras = split(str,"</P>");
paras(end) = [];                % the split left an empty last entry
paras = extractAfter(paras,">") % Drop the "<P ID=n>" from the beginning

Then, numel(paras) will give you the number of </P>.

If you do not have extractFileText, calling string(fileread('file.txt')) should work just fine, too.

In one of the comments, you indicated you also need to count the frequency of words in documents. That is what bagOfWords is for:

tdoc = tokenizedDocument(lower(paras));
bag = bagOfWords(tdoc)
bag = 
bagOfWords with 13 words and 3 documents:
      a   little   black   bird   .   …
      1        1       1      1   1
      1        0       0      1   0
      …

2 Commenti
Mostra NessunoNascondi Nessuno

John il 7 Nov 2017

Apri in MATLAB Online

Thanks. I am stuck running a counter.

for kk = 1:n
  str = ['<p id=',num2str(kk),'>'];
  idx_s = find(strcmp(C,str));
  if ~isempty(idx_s)
      idx_e = idx2(find(idx2>idx_s,1));
      Doc=C(idx_s:idx_e); %May need to remove tags later
      Doc = regexp(Doc,'[a-z0-9\-]+','match');
      Doc = [Doc{:}];
      Unique_Doc_count = arrayfun(@(x) nnz(strcmp(x,Doc)), Unique);
      Unique_Doc_freq=[Unique;Unique_Doc_count];
  end    
end

I want to search if the elements in string array 'Unique' exist in 'Doc'. I got results in 'Unique_Doc_count' as the number of their occurrences but I need just 1 or 0 values (exist) or (not exist). The aim is to loop 'kk' over multiple documents and find the number of documents that contain each word in 'Unique'. Not even number of times the word occurs, but number of documents it appears in.

shilpa patil il 23 Set 2019

Modificato: shilpa patil il 23 Set 2019

how to rewrite the above code for a document image

instead of text file

Accedi per commentare.

Text Extraction and retrieval

0 Commenti
Mostra -2 commenti meno recentiNascondi -2 commenti meno recenti

Risposta accettata

6 Commenti
Mostra 4 commenti meno recentiNascondi 4 commenti meno recenti

Più risposte (2)

6 Commenti
Mostra 4 commenti meno recentiNascondi 4 commenti meno recenti

2 Commenti
Mostra NessunoNascondi Nessuno

Vedere anche

Categorie

Tag

Prodotti

Community Treasure Hunt

Text Extraction and retrieval

0 Commenti Mostra -2 commenti meno recentiNascondi -2 commenti meno recenti

Risposta accettata

6 Commenti Mostra 4 commenti meno recentiNascondi 4 commenti meno recenti

Più risposte (2)

6 Commenti Mostra 4 commenti meno recentiNascondi 4 commenti meno recenti

2 Commenti Mostra NessunoNascondi Nessuno

Vedere anche

Categorie

Tag

Prodotti

Community Treasure Hunt

0 Commenti
Mostra -2 commenti meno recentiNascondi -2 commenti meno recenti

6 Commenti
Mostra 4 commenti meno recentiNascondi 4 commenti meno recenti

6 Commenti
Mostra 4 commenti meno recentiNascondi 4 commenti meno recenti

2 Commenti
Mostra NessunoNascondi Nessuno