<?xml version="1.0" encoding="utf-8"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
	<channel>
		<title>ProZ.com Translation Forums</title>
		<link>http://arm.proz.com/forums/</link>
		<atom:link href="http://arm.proz.com/forums/" rel="self" type="application/rss+xml"/>		<description>Topic: Acronym Extraction/Mining Software</description>
		<language>en-us</language>
		<pubDate>Sat, 08 Aug 2026 19:36:01 +0000</pubDate>
		<lastBuildDate>Sat, 08 Aug 2026 19:36:01 +0000</lastBuildDate>
		<docs>http://www.proz.com/faq</docs>
		<managingEditor>support@proz.com (ProZ.com Support)</managingEditor>
		<webMaster>support@proz.com (ProZ.com Support)</webMaster>
		<item>
			<title>Acronym Extraction/Mining Software | very useful!</title>
			<author>Michael Beijer</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2976713#2976713</link>
			<pubDate>Tue, 08 Nov 2022 18:49:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Michael Beijer&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; very useful!&lt;br/&gt;&lt;br/&gt;[quote]demondragon wrote:&lt;br /&gt;&lt;br /&gt;Hi,&lt;br /&gt;&lt;br /&gt;I realize this is an old question but since no satisfying answer has been provided (the free add-on mentioned above seems to be extinct) I thought I&#039;d chime in with tool I just created.&lt;br /&gt;&lt;br /&gt; [url removed] &lt;br /&gt;&lt;br /&gt;It&#039;s on the web so it should work on any operating system, but it does its acronym and definition &quot;extraction&quot; locally on your computer so you can also be confident that the content of your document won&#039;t end up in someone else&#039;s hands.&lt;br /&gt;&lt;br /&gt;It finds acronyms broadly so there will likely be more than the ones you want but the cost of deleting a false positive (one click) feels justified compared to the impact of potentially missing acronyms in your text.&lt;br /&gt;&lt;br /&gt;It will also attempt to fill the list with each acronym&#039;s definition provided it was written down the first time the acronym is used, like so:&lt;br /&gt;&lt;br /&gt;&quot;A central processing unit (CPU) is...&quot;&lt;br /&gt;&lt;br /&gt;Let me know if it doesn&#039;t help! (And if it&#039;s close to being helpful let me know what&#039;s missing  :) ) [/quote]&lt;br /&gt;&lt;br /&gt;I&#039;ve added a link to your website on my wiki:  [url removed] &lt;br /&gt;&lt;br /&gt;Michael</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | &quot;There&#039;s an app for that&quot; https://listofacronyms.com</title>
			<author>demondragon</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2976691#2976691</link>
			<pubDate>Tue, 08 Nov 2022 15:27:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; demondragon&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; &quot;There&#039;s an app for that&quot; https://listofacronyms.com&lt;br/&gt;&lt;br/&gt;Hi,&lt;br /&gt;&lt;br /&gt;I realize this is an old question but since no satisfying answer has been provided (the free add-on mentioned above seems to be extinct) I thought I&#039;d chime in with tool I just created.&lt;br /&gt;&lt;br /&gt; [url removed] &lt;br /&gt;&lt;br /&gt;It&#039;s on the web so it should work on any operating system, but it does its acronym and definition &quot;extraction&quot; locally on your computer so you can also be confident that the content of your document won&#039;t end up in someone else&#039;s hands.&lt;br /&gt;&lt;br /&gt;It finds acronyms broadly so there will likely be more than the ones you want but the cost of deleting a false positive (one click) feels justified compared to the impact of potentially missing acronyms in your text.&lt;br /&gt;&lt;br /&gt;It will also attempt to fill the list with each acronym&#039;s definition provided it was written down the first time the acronym is used, like so:&lt;br /&gt;&lt;br /&gt;&quot;A central processing unit (CPU) is...&quot;&lt;br /&gt;&lt;br /&gt;Let me know if it doesn&#039;t help! (And if it&#039;s close to being helpful let me know what&#039;s missing  :) )</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | PerfectIt</title>
			<author>Mark</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2588629#2588629</link>
			<pubDate>Thu, 15 Sep 2016 13:37:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Mark&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; PerfectIt&lt;br/&gt;&lt;br/&gt;Another forum user in another thread pointed out this package. They also provide a couple of free add-ons/apps and one of them seems to do this job:&lt;br /&gt;&lt;br /&gt; [url removed] </description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | @Spaddock</title>
			<author>Samuel Murray</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2463976#2463976</link>
			<pubDate>Sun, 30 Aug 2015 09:46:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Samuel Murray&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; @Spaddock&lt;br/&gt;&lt;br/&gt;[quote]Spaddock wrote:&lt;br /&gt;We have not encountered that problem ourselves (because we always start with a proper draft written by a qualified science writer), but if this is a frequent problem, we can easily add this functionality to MAX and let the regex run all across the text and provide best matches together with a likelihood score. [/quote]&lt;br /&gt;&lt;br /&gt;Your tool is designed for documents that have already been &quot;fixed&quot; by an editor (or that were written by authors who did not neglect to write full forms and acronyms together).  The original poster wanted a tool that the editor would use to fix a text written by someone who did neglect it, or that a translator would use to translate such a text.&lt;br /&gt;&lt;br /&gt;[quote]How sure could we be that the acronym is indeed found spelled out &lt;b&gt;somewhere&lt;/b&gt; in the text? [/quote]&lt;br /&gt;&lt;br /&gt;You can&#039;t be sure of that.  That is what software is for.  The software helps prevent the translator/editor from having to manually search the text to see if the author had perhaps spelled out the acronym elsewhere in the text.&lt;br /&gt;&lt;br /&gt;[quote]Is it possible that the author &lt;b&gt;never&lt;/b&gt; mentions it? [/quote]&lt;br /&gt;&lt;br /&gt;Yes, precisely, that is very real risk.  The software would help the translator by letting him know very quickly if the full form likely does not occur elsewhere in the text, thus saving the translator from needlessly going looking for it.&lt;br /&gt;&lt;br /&gt;[quote]Would it, therefore, make sense to add an internet search (based on the content of the text) to see what else it could mean? [/quote]&lt;br /&gt;&lt;br /&gt;That is the usual last resort, yes.  But an internet search can only show you what other authors used the acronym for, and then you have to make an educated guess as to whether your author and that author had used the acronym for the same thing.&lt;br /&gt;&lt;br /&gt;</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | Great! I want it!</title>
			<author>Erik Freitag</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2463970#2463970</link>
			<pubDate>Sun, 30 Aug 2015 09:27:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Erik Freitag&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; Great! I want it!&lt;br/&gt;&lt;br/&gt;[quote]Spaddock wrote:&lt;br /&gt;&lt;br /&gt;We have not encountered that problem ourselves (because we always start with a proper draft written by a qualified science writer), but if this is a frequent problem, we can easily add this functionality to MAX and let the regex run all across the text and provide best matches together with a likelihood score. [/quote]&lt;br /&gt;&lt;br /&gt;That would be excellent, and I&#039;d certainly buy such an app (while its current functionality is of no use to me, or most other translators, I suppose).&lt;br /&gt;&lt;br /&gt;[quote]Spaddock wrote:&lt;br /&gt;How sure could we be that the acronym is indeed found spelled out &lt;b&gt;somewhere&lt;/b&gt; in the text? Is it possible that the author &lt;b&gt;never&lt;/b&gt; mentions it? Would it, therefore, make sense to add an internet search (based on the content of the text) to see what else it could mean?&lt;br /&gt;[/quote]&lt;br /&gt;&lt;br /&gt;In my experience, the chances that an acronym is spelled out somewhere in the text are somewhere around 90%.&lt;br /&gt;&lt;br /&gt;A context based internet search could be a good idea, but for starters, I&#039;ll be happy with the (much more reliable) search within the actual text.&lt;br&gt;&lt;br&gt;[Bearbeitet am 2015-08-30 09:27 GMT]</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | useful and necessary additions</title>
			<author>Michael Beijer</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2463968#2463968</link>
			<pubDate>Sun, 30 Aug 2015 09:16:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Michael Beijer&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; useful and necessary additions&lt;br/&gt;&lt;br/&gt;[quote]Spaddock wrote:&lt;br /&gt;&lt;br /&gt;We have not encountered that problem ourselves (because we always start with a proper draft written by a qualified science writer), but if this is a frequent problem, we can easily add this functionality to MAX and &lt;b&gt;let the regex run all across the text&lt;/b&gt; and provide best matches together with a likelihood score. &lt;br /&gt;&lt;br /&gt;How sure could we be that the acronym is indeed found spelled out &lt;b&gt;somewhere&lt;/b&gt; in the text? Is it possible that the author &lt;b&gt;never&lt;/b&gt; mentions it? Would it, therefore, make sense to add an &lt;b&gt;internet search (based on the content of the text) to see what else it could mean&lt;/b&gt;?&lt;br /&gt;&lt;br /&gt;Thanks! [/quote]&lt;br /&gt;&lt;br /&gt;Hi Spaddock,&lt;br /&gt;&lt;br /&gt;Both very useful and necessary additions. Pity your tool doesn&#039;t run on Windows. Most translators I know run Windows. Myself included.&lt;br /&gt;&lt;br /&gt;Michael</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | MAX 2.0 - this is exactly the kind of feedback we need</title>
			<author>Spaddock</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2463923#2463923</link>
			<pubDate>Sun, 30 Aug 2015 00:33:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Spaddock&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; MAX 2.0 - this is exactly the kind of feedback we need&lt;br/&gt;&lt;br/&gt;We have not encountered that problem ourselves (because we always start with a proper draft written by a qualified science writer), but if this is a frequent problem, we can easily add this functionality to MAX and let the regex run all across the text and provide best matches together with a likelihood score. &lt;br /&gt;&lt;br /&gt;How sure could we be that the acronym is indeed found spelled out &lt;b&gt;somewhere&lt;/b&gt; in the text? Is it possible that the author &lt;b&gt;never&lt;/b&gt; mentions it? Would it, therefore, make sense to add an internet search (based on the content of the text) to see what else it could mean?&lt;br /&gt;&lt;br /&gt;Thanks!</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | MAX isn&#039;t quite what we mean</title>
			<author>Samuel Murray</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2463884#2463884</link>
			<pubDate>Sat, 29 Aug 2015 18:03:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Samuel Murray&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; MAX isn&#039;t quite what we mean&lt;br/&gt;&lt;br/&gt;[quote]Spaddock wrote:&lt;br /&gt;My company has put an OSX App on the App Store that extracts acronyms and their definitions (which are assumed to be found in front of the acronym) from text documents. [/quote]&lt;br /&gt;&lt;br /&gt;Yes, MAX assumes that the acronym is defined in front of the acronym.  In my target language, we sometimes use the acronym first and put the definition in brackets afterwards, but not always.&lt;br /&gt;&lt;br /&gt;Even so, this is not (I think) what is needed by most respondents in this thread.  What we are referring is to a situation in which an acronym is used on its own in the text, without begin defined, but which does have full form somewhere else in the text.&lt;br /&gt;&lt;br /&gt;For example, if the author of the text assumed that his reader knows that ABC means &quot;Apple Bureau for Certification&quot;, he might mention &quot;Apple Bureau for Certification&quot; or even &quot;Apple&#039;s bureau for certification&quot; somewhere in his document while using &quot;ABC&quot; elsewhere in the same document.  What is needed is a program that can guess what &quot;ABC&quot; is the abbreviation of.&lt;br/&gt;&lt;br/&gt;[Edited at 2015-08-29 18:04 GMT]</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | MAX - My Acronym eXtractor - an OSX App that extracts acronyms and their definitions from text files</title>
			<author>Spaddock</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2463847#2463847</link>
			<pubDate>Sat, 29 Aug 2015 15:06:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Spaddock&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; MAX - My Acronym eXtractor - an OSX App that extracts acronyms and their definitions from text files&lt;br/&gt;&lt;br/&gt;My company has put an OSX App on the App Store that extracts acronyms and their definitions (which are assumed to be found in front of the acronym) from text documents.&lt;br /&gt;&lt;br /&gt;You can find it here:&lt;br /&gt; [url removed] &lt;br /&gt;&lt;br /&gt;If you are outside of the countries in which we offer it, send us a message, and we will add your country (we use a default list of countries because we don&#039;t want to worry about paying taxes in countries where we don&#039;t actually have a great demand).&lt;br /&gt;&lt;br /&gt;Here is a link to a video that describes how it works:&lt;br /&gt; [url removed] &lt;br /&gt;&lt;br /&gt;We regularly get requests (mostly from PhD students) who work on PCs to make a PC version, but we don&#039;t have such plans at the moment. Usually, people will have a friend/colleague with a Mac and borrow it to create the first draft of the list of acronyms, and this step alone saves tons of times.&lt;br /&gt;&lt;br /&gt;Of course, MAX is a computer program and not perfect at finding 100% of all acronyms, but it finds the vast majority and saves technical writers tons of times by doing that.&lt;br /&gt;&lt;br /&gt;Hope this helps.</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | Yes</title>
			<author>FarkasAndras</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2438227#2438227</link>
			<pubDate>Fri, 12 Jun 2015 20:46:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; FarkasAndras&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; Yes&lt;br/&gt;&lt;br/&gt;Limiting any such collection effort to specific reputable sources and/or collecting entries by domain would probably be a good idea. Otherwise there may be too much noise and too little signal coming from the resulting termbase.</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | I was thinking of something more targeted myself.</title>
			<author>Mark</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2437996#2437996</link>
			<pubDate>Fri, 12 Jun 2015 10:31:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Mark&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; I was thinking of something more targeted myself.&lt;br/&gt;&lt;br/&gt;[quote]Michael Beijer wrote:&lt;br /&gt;&lt;br /&gt;Ideally, we would take your script, plug it into some kind of (open source) web crawler and scour the internet[/quote]Speaking for myself, I was more interested in looking at the source documents of specific translations rather than creating the ultimate repository of internet acronyms. I’m somewhat sceptical about the idea: there &lt;i&gt;is &lt;/i&gt; a fair bit of nonsense on the web, isn’t there? I wonder what you would really gain from cataloguing it.&lt;br /&gt;&lt;br /&gt;I use those acronym sites very infrequently, imagining that if an acronym is widely used enough to be considered sufficient to express the concept, I’ll be able to find it on my own.</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | &quot;We&quot; is still just me, and is currently on the back burner, but…</title>
			<author>Michael Beijer</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2435771#2435771</link>
			<pubDate>Fri, 05 Jun 2015 14:55:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Michael Beijer&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; &quot;We&quot; is still just me, and is currently on the back burner, but…&lt;br/&gt;&lt;br/&gt;[quote]FarkasAndras wrote:&lt;br /&gt;&lt;br /&gt;[quote]Michael Beijer wrote:&lt;br /&gt;&lt;br /&gt;Ideally, we would take your script, plug it into some kind of (open source) web crawler and scour the internet, and then dump all the data into an open source db (available online, via something like my own fledgling project:  [url removed]  ).&lt;br /&gt;&lt;br /&gt;If the data was available in some form of delimited UTF-8 text format, people could then download it and convert it for use in their own CAT tools.&lt;br /&gt;[/quote]&lt;br /&gt;&lt;br /&gt;Not entirely against the idea. If there is a specific &quot;we&quot; that wants to do this and is willing to put in the time, hit me up via email. [/quote]&lt;br /&gt;&lt;br /&gt;…I&#039;ll drop you a line when I get around to working on the idea a bit more! &lt;br /&gt;&lt;br /&gt;I&#039;m still trying to devise an optimal way to work acronyms and abbreviations into my daily workflow while translating (with CafeTran). I have masses of them, in various formats, but can&#039;t figure out a way so they are available when I need them but don&#039;t clutter up my view when translating.&lt;br /&gt;&lt;br /&gt;I currently have an IntelliWebSearch shortcut set up to simultaneously search several of the leading acronym sites online and I have my massive db as a tab-del glossary in CafeTran.&lt;br /&gt;&lt;br /&gt;I think that a collaboratively maintained mega-list would be a &lt;b&gt;very&lt;/b&gt; valuable resource for us translators.&lt;br /&gt;</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | maybe</title>
			<author>FarkasAndras</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2435170#2435170</link>
			<pubDate>Thu, 04 Jun 2015 10:16:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; FarkasAndras&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; maybe&lt;br/&gt;&lt;br/&gt;[quote]Michael Beijer wrote:&lt;br /&gt;&lt;br /&gt;Ideally, we would take your script, plug it into some kind of (open source) web crawler and scour the internet, and then dump all the data into an open source db (available online, via something like my own fledgling project:  [url removed]  ).&lt;br /&gt;&lt;br /&gt;If the data was available in some form of delimited UTF-8 text format, people could then download it and convert it for use in their own CAT tools.&lt;br /&gt;[/quote]&lt;br /&gt;&lt;br /&gt;Not entirely against the idea. If there is a specific &quot;we&quot; that wants to do this and is willing to put in the time, hit me up via email.</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | No need</title>
			<author>neilmac</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2432654#2432654</link>
			<pubDate>Wed, 27 May 2015 09:56:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; neilmac&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; No need&lt;br/&gt;&lt;br/&gt;Just tell the perpetrators (the people who use acronyms without defining them) that it&#039;s up to them to define the blessed things. In my experience, it never crosses their minds that they might be a mystery to many. &lt;br /&gt;One of my basic conditions for collaboration is the understanding that only the most common and widely understood acronyms (BBC, EU, USA, IMF...) will be translated, while the authors must take responsibility for their more recondite cousins. KWIM?</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | Ideally …</title>
			<author>Michael Beijer</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2432375#2432375</link>
			<pubDate>Tue, 26 May 2015 13:41:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Michael Beijer&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; Ideally …&lt;br/&gt;&lt;br/&gt;[quote]FarkasAndras wrote:&lt;br /&gt;&lt;br /&gt;[quote]Mark Dobson wrote:&lt;br /&gt;&lt;br /&gt;I suppose it could also be done manually with regular expressions, but I’m not really that confident with them myself. If I imagine that I don&#039;t know what HRH stands for, for example:&lt;br /&gt;&lt;br /&gt;H|h\w+\sR|r\w+\sH|h&lt;br /&gt;&lt;br /&gt;I suspect that’s not right, but the idea is that that should match, say:&lt;br /&gt;&lt;br /&gt;her royal highness&lt;br /&gt;hall roof hat&lt;br /&gt;high rumble hip&lt;br /&gt;&lt;br /&gt;And then I could work out the rest on my own. In any case, I’m still surprised not to be able to find something to automate this. Am I missing something?&lt;br /&gt;&lt;br /&gt; [/quote]&lt;br /&gt;&lt;br /&gt;That&#039;s pretty much the basis of how I did it, although of course there is quite a bit more nuance to it than that. I admit that I didn&#039;t do much research to see if there is ready-made open software available for the purpose. If the task is not very complicated, it&#039;s often more convenient to write the code yourself than to try and get someone else&#039;s code working - you often struggle to get it to run, wonder how you&#039;re supposed to use it or whether it&#039;s doing exactly what you want it to.&lt;br /&gt;I only went after acronyms that occur along with the expanded form (I don&#039;t see much of a point in collecting just an acronym that you will have to research from scratch anyway). The easiest way to do it is to find patterns like this: L1\w+ L2\w+ L3\w+ \(L1L2L3\). I collected acronyms in multiple languages based on aligned texts so there is some wizardry in trying to make sure that they are paired up correctly and there is a lot of fiddling in covering various kinds of unusual cases (E.g. CITES is the Convention on International Trade in Endangered Species, which won&#039;t be picked up by a primitive pattern search that doesn&#039;t know that &quot;on&quot; and &quot;in&quot; are filler words. Even worse, if your acronym is the framework for international bartending standards or the Organisation Of European Fortune Tellers, a primitive algorithm will chop the first word off. Ask how I know.)&lt;br /&gt;I had a quick look at the linked MS paper, and it looks like they went a LOT deeper down this rabbit hole with AcroMiner than I did. It&#039;s a shame they didn&#039;t publish the code. I&#039;m not even sure there&#039;s any point in publishing what looks like it was intended as a scientific paper and then holding back the actual goods.&lt;br /&gt;&lt;br /&gt;Not sure if I want to share my script... It&#039;s pretty rough and designed to work with my specific EU files. It could be polished up a little and upgraded to work with other files, but that would be a fair bit of work and it would still be inferior to better researched software like acrominer. If no other (better) tool is available online I might be persuaded to do it.&lt;br/&gt;&lt;br/&gt;[Edited at 2015-05-22 14:46 GMT] [/quote]&lt;br /&gt;&lt;br /&gt;Ideally, we would take your script, plug it into some kind of (open source) web crawler and scour the internet, and then dump all the data into an open source db (available online, via something like my own fledgling project:  [url removed]  ).&lt;br /&gt;&lt;br /&gt;If the data was available in some form of delimited UTF-8 text format, people could then download it and convert it for use in their own CAT tools.&lt;br/&gt;&lt;br/&gt;[Edited at 2015-05-26 14:34 GMT]</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | regex</title>
			<author>FarkasAndras</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2431390#2431390</link>
			<pubDate>Fri, 22 May 2015 14:41:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; FarkasAndras&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; regex&lt;br/&gt;&lt;br/&gt;[quote]Mark Dobson wrote:&lt;br /&gt;&lt;br /&gt;I suppose it could also be done manually with regular expressions, but I’m not really that confident with them myself. If I imagine that I don&#039;t know what HRH stands for, for example:&lt;br /&gt;&lt;br /&gt;H|h\w+\sR|r\w+\sH|h&lt;br /&gt;&lt;br /&gt;I suspect that’s not right, but the idea is that that should match, say:&lt;br /&gt;&lt;br /&gt;her royal highness&lt;br /&gt;hall roof hat&lt;br /&gt;high rumble hip&lt;br /&gt;&lt;br /&gt;And then I could work out the rest on my own. In any case, I’m still surprised not to be able to find something to automate this. Am I missing something?&lt;br /&gt;&lt;br /&gt; [/quote]&lt;br /&gt;&lt;br /&gt;That&#039;s pretty much the basis of how I did it, although of course there is quite a bit more nuance to it than that. I admit that I didn&#039;t do much research to see if there is ready-made open software available for the purpose. If the task is not very complicated, it&#039;s often more convenient to write the code yourself than to try and get someone else&#039;s code working - you often struggle to get it to run, wonder how you&#039;re supposed to use it or whether it&#039;s doing exactly what you want it to.&lt;br /&gt;I only went after acronyms that occur along with the expanded form (I don&#039;t see much of a point in collecting just an acronym that you will have to research from scratch anyway). The easiest way to do it is to find patterns like this: L1\w+ L2\w+ L3\w+ \(L1L2L3\). I collected acronyms in multiple languages based on aligned texts so there is some wizardry in trying to make sure that they are paired up correctly and there is a lot of fiddling in covering various kinds of unusual cases (E.g. CITES is the Convention on International Trade in Endangered Species, which won&#039;t be picked up by a primitive pattern search that doesn&#039;t know that &quot;on&quot; and &quot;in&quot; are filler words. Even worse, if your acronym is the framework for international bartending standards or the Organisation Of European Fortune Tellers, a primitive algorithm will chop the first word off. Ask how I know.)&lt;br /&gt;I had a quick look at the linked MS paper, and it looks like they went a LOT deeper down this rabbit hole with AcroMiner than I did. It&#039;s a shame they didn&#039;t publish the code. I&#039;m not even sure there&#039;s any point in publishing what looks like it was intended as a scientific paper and then holding back the actual goods.&lt;br /&gt;&lt;br /&gt;Not sure if I want to share my script... It&#039;s pretty rough and designed to work with my specific EU files. It could be polished up a little and upgraded to work with other files, but that would be a fair bit of work and it would still be inferior to better researched software like acrominer. If no other (better) tool is available online I might be persuaded to do it.&lt;br/&gt;&lt;br/&gt;[Edited at 2015-05-22 14:46 GMT]</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | Thank you, both</title>
			<author>Mark</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2431172#2431172</link>
			<pubDate>Fri, 22 May 2015 07:40:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Mark&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; Thank you, both&lt;br/&gt;&lt;br/&gt;I might have a fiddle with CafeTran at home then (I don’t imagine my employers would bother installing it on the system for the one function; they’re bound to tell me I should be doing it with UltraEdit). CafeTran and CafeTran Espresso are, I gather, different ways of saying the same thing?&lt;br /&gt;&lt;br /&gt;Since András is selling the glossaries and suggests that people who want &quot;to extract terminological data from [large amounts of text] to create specialized termbases/glossaries like these […] get in touch&quot;, I imagine that he’s decided to keep his methods for himself. Perhaps I’ll contact him to make sure though; it seems to me he could sell his work in another way if he chose to.</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | online/local acronym-mining tool w/ expanded form finder</title>
			<author>Michael Beijer</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2430767#2430767</link>
			<pubDate>Thu, 21 May 2015 11:09:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Michael Beijer&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; online/local acronym-mining tool w/ expanded form finder&lt;br/&gt;&lt;br/&gt;An acronym-mining tool would indeed be very useful, especially if it was also able to find potential expanded forms. It would also be great if it could be let loose on the internet.&lt;br /&gt;&lt;br /&gt;related stuff:&lt;br /&gt;&lt;br /&gt;Havbe you seen András Farkas’s big EU acronym collection? &lt;br /&gt;&lt;br /&gt;@  [url removed] &lt;br /&gt;&lt;br /&gt;[quote]&quot;EU acronym collection - NOW AVAILABLE!&lt;br /&gt;&lt;br /&gt;Similarly to the glossary, the acronyms were also harvested from aligned document sets &lt;b&gt;using custom software tools made for this purpose&lt;/b&gt;, with a limited amount of manual correction. The acronym collection is available as a bilingual or multilingual glossary. In bilingual versions, each entry contains four fields: the acronym in language 1, the full expression in language 1, the acronym in language 2 and the full expression in language 2. E.g. ETO / European telecommunications office / BET / Bureau européen des télécommunications. In some entries, certain fields (full expression in languages other than English) are empty. In most cases, this is because the language in question uses the English acronym and thus the letters of the acronym don&#039;t match the full form, which prevents automated recognition. Every English acronym is listed along with the corresponding full English expression, and detailed statistics on other languages are available on request.&lt;br /&gt;The acronym collection covers a vast range of areas, with entries ranging from the British Aluminium Foil Rollers Association (BAFRA) to the International Plant Protection Convention (IPPC, French: CIPV, Convention internationale pour la protection des végétaux). There are about 8,000 entries in all, with the potential to save you untold hours of tedious research. The acronym collection covers all EU languages except Croatian and Irish. The number of entries depends on the language pair requested.&lt;br /&gt;A sample is available here (xls). The sample file contains the English, French and German versions of all the acronyms that start with the letter A.&lt;br /&gt;Formats: tab delimited txt, xls and tmx. Other formats (tbx, xml etc.) available on request. I recommend using this data as a termbase, not a TM (i.e. import it into MultiTerm or the terminology module of your CAT of choice). If your terminology software can&#039;t handle synonyms (e.g. two English columns and two French columns in the same table), let me know and I will create a special two-column version that allows both the acronyms and the full forms to be all imported into the same database.&lt;br /&gt;Price: EUR 25 for a bilingual glossary, plus EUR 10 for each additional language.&quot;[/quote]&lt;br /&gt;&lt;br /&gt;You might want to ask him about those &quot;custom software tools made for this purpose&quot;.&lt;br /&gt;&lt;br /&gt;I am also working on my own collection (in my limited spare time):  [url removed] &lt;br /&gt;However, my collection methods are much more lo-fi: I simply scour the internet in search of lists of acronyms, and also extract content from existing collections using scraping tools like HTTrack.&lt;br /&gt;&lt;br /&gt;Another related idea is to use IntelliWebSearch to search sites like  [url removed]  with a Windows shortcut, for when you come across a pesky one while translating.&lt;br /&gt;&lt;br /&gt;As a happy CafeTran user, I thought I&#039;d also mention the automatic abbreviation extractor, under:&lt;br /&gt;&lt;br /&gt;&lt;b&gt;Tools &gt; Abbreviations &gt; Scan Project for abbreviations&lt;/b&gt;&lt;br /&gt;</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software | CafeTran can extract acronyms</title>
			<author>Igor Kmitowski</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2430751#2430751</link>
			<pubDate>Thu, 21 May 2015 10:41:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Igor Kmitowski&lt;br/&gt;&lt;b&gt;Post title:&lt;/b&gt; CafeTran can extract acronyms&lt;br/&gt;&lt;br/&gt;Hi Mark,&lt;br /&gt;&lt;br /&gt;The feature is available in the free version of CafeTran Espresso.&lt;br /&gt;&lt;br /&gt;1. Create a Project with your source document(s).&lt;br /&gt;2. Go to Edit &gt; Find... panel.&lt;br /&gt;3. Select Project Source Segments scope.&lt;br /&gt;4. Select Regular expression box and Extract reg. exp. results box.&lt;br /&gt;5. Type your regular expression and click Find.&lt;br /&gt;&lt;br /&gt;CafeTran will list the results in one simple text column that you can save or copy.&lt;br /&gt;&lt;br /&gt;Igor</description>
		</item>
		<item>
			<title>Acronym Extraction/Mining Software</title>
			<author>Mark</author>
			<category>Software applications</category>
			<link>http://arm.proz.com/post/2430728#2430728</link>
			<pubDate>Thu, 21 May 2015 09:44:00 +0000</pubDate>
			<description>&lt;b&gt;Forum:&lt;/b&gt; Software applications&lt;br/&gt;&lt;b&gt;Topic:&lt;/b&gt; Acronym Extraction/Mining Software&lt;br/&gt;&lt;b&gt;Poster:&lt;/b&gt; Mark&lt;br/&gt;&lt;br/&gt;Hi,&lt;br /&gt;&lt;br /&gt;I had a thought a while back, about acronyms, and I was just speaking to a colleague of mine who had the same thought independently. How come there isn’t an application (that we know of, at least) that mines acronyms and their potential expanded forms from documents? We can’t be the only people who translate pages full of impenetrable acronyms that you suspect are hiding in plain sight in the text, can we?&lt;br /&gt;&lt;br /&gt;It’s not hard to find proposals on how to do this online, with algorithms an suchlike, but the actual applications seem more elusive:&lt;br /&gt;&lt;br /&gt; [url removed] &lt;br /&gt; [url removed] &lt;br /&gt;&lt;br /&gt;I found a Word add-on that lets you simply extract the acronyms into a table for you to define yourself at a later date, but that seems like a job half done to me.&lt;br /&gt;&lt;br /&gt;I suppose it could also be done manually with regular expressions, but I’m not really that confident with them myself. If I imagine that I don&#039;t know what HRH stands for, for example:&lt;br /&gt;&lt;br /&gt;H|h\w+\sR|r\w+\sH|h&lt;br /&gt;&lt;br /&gt;I suspect that’s not right, but the idea is that that should match, say:&lt;br /&gt;&lt;br /&gt;her royal highness&lt;br /&gt;hall roof hat&lt;br /&gt;high rumble hip&lt;br /&gt;&lt;br /&gt;And then I could work out the rest on my own. In any case, I’m still surprised not to be able to find something to automate this. Am I missing something?&lt;br /&gt;&lt;br /&gt;</description>
		</item>
	</channel>
</rss>